Skip to main content

Add opt-in option to automatically clean up failed worker pods

Currently, Spacelift WorkerPools retain failed pods indefinitely to allow for debugging. While this is useful for short-term analysis, in certain scenarios (e.g., when a stack’s Project root is removed in Git but drift detection remains active), these failed pods can accumulate rapidly and exhaust the worker pool.

Proposed Feature:
Introduce an optional configuration (e.g., keepFailedPods=false or a TTL-based cleanup) that allows customers to opt-in to automatic cleanup of failed pods, similar to the existing keepSuccessfulPods flag. This would let customers balance between log retention and cluster stability.Value:

  • Prevents runaway resource consumption when jobs repeatedly fail at the workspace preparation stage.

  • Reduces operational burden of manually cleaning up failed pods.

  • Maintains flexibility by keeping the current default (retain failed pods), but provides an escape hatch for customers prioritizing stability.

Status: ✅ Completed2 comments

Log in to comment and vote

Comments2

  • Jonah Kowall

    Team•

    Aug 17

    This is available now. Pod retention on Kubernetes workers was reworked in workerpool-controller v0.0.25 / Helm chart v0.46.0, and it covers both shapes you proposed, count-based and TTL-based:

    spec:

    failedPodsHistoryLimit: 0 # remove failed pods immediately

    # or keep a bounded window:

    failedPodsHistoryLimit: 10

    failedPodsHistoryTTL: "72h" # and/or drop anything older than 72h

    Defaults are unchanged in spirit: failed pods keep the 5 most recent (failedPodsHistoryLimit: 5), so debugging still works out of the box, and you opt in to more aggressive cleanup. Pods are removed if they exceed either limit, running pods are never affected, and cleanup is managed at the WorkerPool level across every worker in the pool. successfulPodsHistoryLimit / successfulPodsHistoryTTL are the matching pair for successful pods; note that keepSuccessfulPods was deprecated and removed in the same release, so set successfulPodsHistoryLimit to a positive value if you were relying on it.


    Full reference: https://docs.spacelift.io/concepts/worker-pools/kubernetes-workers


    One thing worth separating out: the scenario you hit, a stack whose project root was removed in Git while drift detection kept running, means those pods are a symptom of runs that arguably shouldn't be firing at all. Cleanup stops the pool from filling up, but if that's still happening to you, tell us and we'll look at the drift-detection side directly.

  • Black Breeze

    •

    Aug 25, 2025

    Dom, one of our product engineers will be in touch with you directly to understand this failure mode. Let’s figure out the best approach together. Thanks for escalating 🙇🏼‍♂️