Failure Grouping & Deduplication in Stack Runs
Problem: We regularly end up with hundreds of failed stacks due to unpinned dependencies and provider drift. Today, it’s extremely manual to determine which failures are identical vs. which are unique across environments, components, and regions.
Request:
Add failure grouping in the UI or via a downloadable report, ideally based on stack output. Group failures by exception type or error pattern so teams can quickly identify and resolve the root causes without clicking into every individual stack.
Impact: Helps large orgs triage & remediate faster, saves engineering time, and enables actual prioritization of real vs. noisy issues.
Log in to comment and vote
Comments2
Jonah Kowall
Aug 13
Hi, and sorry for the slow follow-up.
We are not planning to build native failure grouping into the runs view. Error clustering is something observability tools already do well, so we give you the plumbing instead: notification policies can stream every failure to a webhook or Datadog, and the GraphQL API exposes run outcomes so you can group them wherever your team already handles incidents.
On the cause you mentioned, unpinned dependencies and provider drift, these will do more than better triage:
Commit
.terraform.lock.hclso provider versions stop moving under youPin Terraform/OpenTofu versions per stack
Use a plan policy to reject unpinned dependencies at review instead of at deploy
If you want the in-product views to handle this better, that is being tracked here: https://feedback.spacelift.io/p/dashboard-filters-for-drift-and-better-failure-semantics. Worth adding your vote and any details there please.
Natalia Gazda
Aug 5, 2025
Hey Harrison!
Thanks for sharing this detailed feature request.I totally understand how painful it is to manually click through hundreds of failed stacks to identify patterns.
Here are some approaches that might help in the meantime:
1. Use notification policies with webhooks to stream failure data to an external observability platform (Datadog, Splunk, etc.) where you can perform pattern analysis and grouping
2. Leverage our GraphQL API to pull run failure data programmatically and build custom reporting/grouping logic
3. Consider stack naming conventions or labels that help identify related stacks when failures occur (e.g., labeling by component, region, or dependency type)
4. Enable drift detection to catch issues proactively before they cascade into deployment failures
For the unpinned dependencies issue specifically, you might also want to look at using Spacelift policies to enforce dependency pinning during the plan phase. Feel free to reach out to our support team, if help with setting up anything was needed.
I've captured your feedback for future consideration. A bit more discovery questions - are these failures typically happening during regular deployments or during drift detection runs? And roughly how often are you dealing with these mass failure events?