Hey folks! Quick question to better understand the use case here: Have you considered using our Kubernetes-based worker pools? They handle this exact autoscaling pattern really well - avoiding the scale-down/scale-up thrashing you're experiencing.
Is there something specific preventing you from using k8s workers? (No k8s environment, company constraints, specific EC2 requirements, etc?) Understanding what's keeping you on EC2 will help us figure out the best path forward!
It's a pretty common scenario to have a job be triggered not immediately after one that just completed, but maybe seconds, or a minute later. For example, a manual retry on a failed run. If you're "unlucky", the autoscaler's next evaluation will land between the first job finishing and the second starting. This means that the autoscaler will scale the ASG down just before the to-be-killed worker is needed again. When the second job kicks off, it takes at least one autoscaler cycle to spin up a new worker. However, that may be compounded by https://github.com/spacelift-io/terraform-aws-spacelift-workerpool-on-ec2/issues/84, where it may take even longer to have a "ready" worker (since the autoscaler won't add a new worker until the one that's "draining" falls off).
I'd propose a new feature, similar to $AUTOSCALING_MAX_KILL where the autoscaler will delay terminating a worker for a configurable amount of time, effectively slowing scale-down events.
I'm not sure exactly how to approach this, but I have a few ideas:
Maybe there's something in the worker's metadata or the queue metadata that says how long it's been since a worker was busy?
Maybe idle workers can be "marked" for termination, but not executed until the next cycle? Maybe the autoscaler can tag the instance with a pending_termination tag. Then, on the next run, it will either terminate the instance if that tag already exists or remove the tag if it's busy again. That way, any termination will always take 2 cycles.
The latter approach seems relatively simple and elegant, TBH.
Log in to comment and vote
Comments3
Black Breeze
May 30, 2025
Hey folks! Quick question to better understand the use case here: Have you considered using our Kubernetes-based worker pools? They handle this exact autoscaling pattern really well - avoiding the scale-down/scale-up thrashing you're experiencing.
Here's our documentation: https://docs.spacelift.io/concepts/worker-pools/kubernetes-workers
Is there something specific preventing you from using k8s workers? (No k8s environment, company constraints, specific EC2 requirements, etc?) Understanding what's keeping you on EC2 will help us figure out the best path forward!
Beige Antelope
May 30, 2025
We don’t (really) use k8s in our company.
Azure Chipmunk
May 30, 2025
Related:
https://github.com/spacelift-io/ec2-workerpool-autoscaler/issues/25
To quote: