Skip to main content

Delay between scale down events in AWS autoscaler

The terraform-aws-ec2-autoscaler should have a minimum (configurable) delay between scale out and scale in for an individual worker.

Workaround
Problem
Status: ✅ Completed3 comments

Log in to comment and vote

Comments3

  • Black Breeze

    •

    May 30, 2025

    Hey folks! Quick question to better understand the use case here: Have you considered using our Kubernetes-based worker pools? They handle this exact autoscaling pattern really well - avoiding the scale-down/scale-up thrashing you're experiencing.

    Here's our documentation: https://docs.spacelift.io/concepts/worker-pools/kubernetes-workers

    Is there something specific preventing you from using k8s workers? (No k8s environment, company constraints, specific EC2 requirements, etc?) Understanding what's keeping you on EC2 will help us figure out the best path forward!

    • Beige Antelope

      •

      May 30, 2025

      We don’t (really) use k8s in our company.

  • Azure Chipmunk

    •

    May 30, 2025

    Related:

    https://github.com/spacelift-io/ec2-workerpool-autoscaler/issues/25

    To quote:

    It's a pretty common scenario to have a job be triggered not immediately after one that just completed, but maybe seconds, or a minute later. For example, a manual retry on a failed run. If you're "unlucky", the autoscaler's next evaluation will land between the first job finishing and the second starting. This means that the autoscaler will scale the ASG down just before the to-be-killed worker is needed again. When the second job kicks off, it takes at least one autoscaler cycle to spin up a new worker. However, that may be compounded by https://github.com/spacelift-io/terraform-aws-spacelift-workerpool-on-ec2/issues/84, where it may take even longer to have a "ready" worker (since the autoscaler won't add a new worker until the one that's "draining" falls off).

    I'd propose a new feature, similar to $AUTOSCALING_MAX_KILL where the autoscaler will delay terminating a worker for a configurable amount of time, effectively slowing scale-down events.

    I'm not sure exactly how to approach this, but I have a few ideas:

    • Maybe there's something in the worker's metadata or the queue metadata that says how long it's been since a worker was busy?

    • Maybe idle workers can be "marked" for termination, but not executed until the next cycle? Maybe the autoscaler can tag the instance with a pending_termination tag. Then, on the next run, it will either terminate the instance if that tag already exists or remove the tag if it's busy again. That way, any termination will always take 2 cycles.

    The latter approach seems relatively simple and elegant, TBH.