Skip to main content

Failure Grouping & Deduplication in Stack Runs

Problem: We regularly end up with hundreds of failed stacks due to unpinned dependencies and provider drift. Today, it’s extremely manual to determine which failures are identical vs. which are unique across environments, components, and regions.
Request:
Add failure grouping in the UI or via a downloadable report, ideally based on stack output. Group failures by exception type or error pattern so teams can quickly identify and resolve the root causes without clicking into every individual stack.
Impact: Helps large orgs triage & remediate faster, saves engineering time, and enables actual prioritization of real vs. noisy issues.

Status: ❌ Rejected2 comments

Log in to comment and vote

Comments2

  • Jonah Kowall

    Team•

    Aug 13

    Hi, and sorry for the slow follow-up.

    We are not planning to build native failure grouping into the runs view. Error clustering is something observability tools already do well, so we give you the plumbing instead: notification policies can stream every failure to a webhook or Datadog, and the GraphQL API exposes run outcomes so you can group them wherever your team already handles incidents.


    On the cause you mentioned, unpinned dependencies and provider drift, these will do more than better triage:

    • Commit .terraform.lock.hcl so provider versions stop moving under you

    • Pin Terraform/OpenTofu versions per stack

    • Use a plan policy to reject unpinned dependencies at review instead of at deploy


    If you want the in-product views to handle this better, that is being tracked here: https://feedback.spacelift.io/p/dashboard-filters-for-drift-and-better-failure-semantics. Worth adding your vote and any details there please.

  • Natalia Gazda

    Team•

    Aug 5, 2025

    Hey Harrison!

    Thanks for sharing this detailed feature request.I totally understand how painful it is to manually click through hundreds of failed stacks to identify patterns.

    Here are some approaches that might help in the meantime:

    1. Use notification policies with webhooks to stream failure data to an external observability platform (Datadog, Splunk, etc.) where you can perform pattern analysis and grouping
    2. Leverage our GraphQL API to pull run failure data programmatically and build custom reporting/grouping logic
    3. Consider stack naming conventions or labels that help identify related stacks when failures occur (e.g., labeling by component, region, or dependency type)
    4. Enable drift detection to catch issues proactively before they cascade into deployment failures

    For the unpinned dependencies issue specifically, you might also want to look at using Spacelift policies to enforce dependency pinning during the plan phase. Feel free to reach out to our support team, if help with setting up anything was needed.

    I've captured your feedback for future consideration. A bit more discovery questions - are these failures typically happening during regular deployments or during drift detection runs? And roughly how often are you dealing with these mass failure events?