Skip to main content

Allow replan in failed plans

Sometimes a stack can have some failing resources, but at the same time needs some resources to be applied.

In a normal non failing stack, you can use the replan to apply those changes.

But, if the first plan fails, there is no way to do a replan targeting only some the non failing resources, as it would be possible to do in local Terraform execution.

Workaround
Problem
Status: ⬆️ Gathering votes3 comments

Log in to comment and vote

Comments3

  • Natalia Gazda

    Team•

    Jul 19, 2025

    Thanks for sharing this request. To better understand the problem and find the right solution, I'd like to explore a few aspects:

    Understanding the Pattern

    • How frequently do you encounter plans that fail with "some bad resources"? Is this a daily occurrence or more occasional?
    • What typically causes these resource failures - permission issues, API changes, drift, or something else?
    • When you mention targeting "a lot" of resources, roughly how many are we talking about?

    Business Impact

    • What's blocked while you work around these failures? Are critical deployments delayed?
    • How much time does the current task-based workaround typically take?

    Root Cause

    • Have you considered splitting these into separate stacks to isolate failing components?
    • Is there a specific architectural reason these resources need to remain in the same stack?

    Understanding whether this is about handling temporary failures vs. persistent infrastructure issues will help us design the right solution - whether that's better targeting UI or addressing the root causes of mixed-health stacks.

    • Lime Fork

      •

      Aug 11, 2025

      Hi @Natalia Gazda I have been away for a while, I will answer to your questions during this week.

    • Lime Fork

      •

      Aug 22, 2025

      Understanding the Pattern

      • How frequently do you encounter plans that fail with "some bad resources"? Is this a daily occurrence or more occasional?

        • It is rare, but we worry this may happen in a critical moment (e.g. incident resolution)

      • What typically causes these resource failures - permission issues, API changes, drift, or something else?

        • Drift, mostly. There is something here related with how we do Terraform, where in-house modules are not versioned and are part of the monorepo. Very widely use modules can sometimes cause problems in stacks that have drifted and that have been run in a while.

      • When you mention targeting "a lot" of resources, roughly how many are we talking about?

        • From tens to hundreds. Most of the times being able to exclude few resources would be enough

      Business Impact

      • What's blocked while you work around these failures? Are critical deployments delayed?

        • Not at the moment, just affecting development velocity. The worry is that this would happen in a critical moment.

      • How much time does the current task-based workaround typically take?

        • Difficult to say, the few times this has happened it has been from minutes, to a couple of days, depending urgency and cause for the error.

      Root Cause

      • Have you considered splitting these into separate stacks to isolate failing components?

      • Is there a specific architectural reason these resources need to remain in the same stack?


        I feel response to these two are not very relevant in this case, as this does not happen with specific resources. Though it is relevant what I mentioned about non versioned modules in a previous response.



      Btw @Natalia Gazda I guess this would belong more to another feature request, but I the dream would be to be able to select resources from the IaC Management tab and trigger a targeted run from there (without needing to run a complete plan first).

      Let me know if this is helpful or there are other questions or details I could give.

      Thanks!