Skip to main content

Introduce keepFailedPods to cleanup failed pods

We're using the spacelift-workerpool-controller Helm chart and managing our WorkerPools via the provided CRDs. We've noticed that while the keepSuccessfulPods option is available to manage successful pod cleanup, there doesn't seem to be a similar option for failed pods.

We understand that keeping failed pods is useful for debugging purposes. However, over time this can clutter the namespace and create additional maintenance overhead. While we could implement our own cleanup mechanism, we wanted to reach out and ask if there's an existing way to handle this - or if you'd consider introducing a keepFailedPods option to provide more flexibility.

Thanks in advance for the help!

Status: ❌ Rejected8 comments

Log in to comment and vote

Comments8

  • Aquamarine Butter

    •

    Jun 30, 2025

    Hello there, I’m one of the engineers that was in charge of the workerpool controller implementation. I got pulled here into this thread by one of my colleague, and I would like to clarify one thing.

    We decided not to care about failing pods on purpose. The reason is that Kubernetes does have a built-in garbage collector mechanism for failed pods.

    So I’m curious to know why does the built-in GC does not already take care of cleaning up those pods for you. Maybe it’s just a matter of just adjusting the terminated-pod-gc-thresholdsetting?

    Possibly there is another lifecycle quirk that I’m struggling to understand, and maybe the GC does not actually catch those pods because of a specific reason. So would you mind checking if those pods are indeed in a Failed state in the status?

    We’ll not implement our own failed pod eviction logic in the controller, because that’s not quite really the responsibility of the controller.

    On the other hand, there could be an issue in the controller that makes your pod in a specific state that prevents them to be caught by the GC. If yes, that would be great if you could give us a reproducible case so we could find a way to fix that on our side.

    Thanks for raising that anyway!

  • Black Breeze

    •

    May 27, 2025

    Thanks for the feedback! This makes sense given we already have `keepSuccessfulPods`. Before we dive into implementation, I'd love to understand the problem a bit better: - How many failed pods typically accumulate in your namespace over a week or month? - What specific maintenance overhead are you experiencing? (storage pressure, visual clutter, hitting resource limits, etc.) - When pods do fail, how often do you actually need to inspect them for debugging vs. just wanting them cleaned up? - Are there particular types of failures where you'd want to keep the pods longer for investigation? Understanding the real pain points will help us design the right solution!