Add opt-in option to automatically clean up failed worker pods
Currently, Spacelift WorkerPools retain failed pods indefinitely to allow for debugging. While this is useful for short-term analysis, in certain scenarios (e.g., when a stack’s Project root is removed in Git but drift detection remains active), these failed pods can accumulate rapidly and exhaust the worker pool.
Proposed Feature:
Introduce an optional configuration (e.g., keepFailedPods=false or a TTL-based cleanup) that allows customers to opt-in to automatic cleanup of failed pods, similar to the existing keepSuccessfulPods flag. This would let customers balance between log retention and cluster stability.Value:
Prevents runaway resource consumption when jobs repeatedly fail at the workspace preparation stage.
Reduces operational burden of manually cleaning up failed pods.
Maintains flexibility by keeping the current default (retain failed pods), but provides an escape hatch for customers prioritizing stability.
Log in to comment and vote
Comments2
Jonah Kowall
Aug 17
This is available now. Pod retention on Kubernetes workers was reworked in workerpool-controller v0.0.25 / Helm chart v0.46.0, and it covers both shapes you proposed, count-based and TTL-based:
spec:failedPodsHistoryLimit: 0 # remove failed pods immediately# or keep a bounded window:failedPodsHistoryLimit: 10failedPodsHistoryTTL: "72h" # and/or drop anything older than 72hDefaults are unchanged in spirit: failed pods keep the 5 most recent (failedPodsHistoryLimit: 5), so debugging still works out of the box, and you opt in to more aggressive cleanup. Pods are removed if they exceed either limit, running pods are never affected, and cleanup is managed at the WorkerPool level across every worker in the pool. successfulPodsHistoryLimit / successfulPodsHistoryTTL are the matching pair for successful pods; note that keepSuccessfulPods was deprecated and removed in the same release, so set successfulPodsHistoryLimit to a positive value if you were relying on it.
Full reference: https://docs.spacelift.io/concepts/worker-pools/kubernetes-workers
One thing worth separating out: the scenario you hit, a stack whose project root was removed in Git while drift detection kept running, means those pods are a symptom of runs that arguably shouldn't be firing at all. Cleanup stops the pool from filling up, but if that's still happening to you, tell us and we'll look at the drift-detection side directly.
Black Breeze
Aug 25, 2025
Dom, one of our product engineers will be in touch with you directly to understand this failure mode. Let’s figure out the best approach together. Thanks for escalating 🙇🏼♂️