Improved Visibility into Worker Billing (P95 Model)
We’d like to request improvements to the billing visibility, specifically around how worker usage is calculated and billed under the P95 model.
Background:
Currently, the P95 billing model is a bit tricky to fully grasp from the UI, especially when trying to estimate how much we'll be billed or whether we’re within our buffer. We asked a few questions and got helpful answers:
Example:
If we occasionally exceed the limit (e.g., 6 workers for a short period), the P95 might still be within our contracted limit (e.g., 5 workers), meaning no extra charges.
However, if we go far over (e.g., 30 workers for the same short time), this could push the P95 up and cause overages.
So while both duration and degree of overage matter, it’s currently hard to know what our actual risk or buffer is.
Feature Request:
We’d love to see better visibility in the UI and/or via metrics endpoints around this billing model, specifically:
Current P95 usage: Where we stand right now for the billing cycle.
Projected billing impact: A clear estimate of our current or projected overages, if any.
Buffer insights: An answer to the question: “How much buffer do we have left this month?”
Simulation tool: Something that helps us simulate the impact of increasing worker count temporarily (e.g., during migrations between pools) so we can make informed decisions.
Why it matters:
This will help us:
Plan operations like migrations without unintentional overages.
Understand when it's safe to scale up temporarily.
Avoid surprises in billing.
Thanks for your consideration!
Log in to comment and vote
Comments18
Oct 6
PinnedDear Upvoters, this feature is now Live, in your Org settings, together with notifications on reaching 80% and 100% of your burst limit. Looking forward to your feedback!
Azure Chipmunk
Oct 7
It looks great so far! Much clearer.
Black Breeze
May 10, 2025
Would seeing how much burst budget you have left—i.e., how long you can exceed your committed concurrency without increasing your p95—help you confidently plan migrations or scale events? If not, what else would you need to feel safe?
Copper Pretzel
Jul 18, 2025
In addition to the other ideas shared in this conversation, which in an internal discussion all seem favorable to aid us in worker management, we’d like to also propose the potential for a worker policy. If a policy could block workers if it’d take us over budget with the ability to opt-in bypass the policy if we decide the overage is worth it that’d help safeguard our management as well.
Another perspective on the metrics aspect of this topic is to allow us to use the burst budget as a way to request budget for more workers. Running hot a handful of times a month gives us mainly the ability to work within our maintenance windows. If we can calculate that we’re running out of burst window in relation to our maintenance schedule it gives us a solid metric to request a budget review. Otherwise our main argument is engineer wait time and the worker cost doesn’t really scale well to engineer wait time even if they did care about that.
Cyan Crab
May 14, 2025
@Marcin Wyszynski thank you for being responsive to this customer feedback. My team has also been experiencing a lot of frustration around the billing model for private worker pools, and we're thinking about some solutions for this.
I think the idea of a "burst budget" is an interesting one. What about an even simpler suggestion though - show the billing cycle in the web UI.
The billing page (/settings/billing) just shows general info on the plan. The "plan details" button just links to https://spacelift.io/pricing and doesn't actually show specific plan details. So this page is not very useful.
The usage page just shows the last 30 days of usage, which is not necessarily the billing cycle. Similarly, the "Export CSV" button downloads a CSV of usage data for the last 30 days, not for the billing cycle. So if we get charged for overages, we can't actually get the data that you used to calculate those overages.
So my suggestion is for the billing and usage pages to indicate when the billing cycle is (meaning which day of the month the billing cycle ends on), and should give customers the ability to download usage data for the billing cycle, not just the last 30 days.
Indigo Spacelab
May 26, 2025
We’re in the same boat. We typically operate up to 5 workers, but there are certain refactors or initiatives that require updates to multiple stacks over a short period of time, therefore prompting us to allocate additional workers for that time frame. But with the current reporting in Spacelift and the percentile system, it’s really difficult to estimate how many workers we can use, and for what amount of time, before we breach our contract limits. Overages in those situations could end up costly if we estimate the number incorrectly.
Moccasin Wobble
Sep 3, 2025
I think it would be great to have this feature. Something that would even take it to the next level is if it could be integrated with Spacelift’s kubernetes self-hosted runner autoscaler. That way, you could specify your own min and max concurrency and level of full time workers you have allotted for the month. Operating at a level less than your agreed number of workers per month would “bank up” minutes. The autoscaler could look at how many minutes you have banked up before bursting. If you don’t have any in the bank, it can’t burst and only scales up to the number of allotted full time workers you have for the month (or scales down to that if you run out of burst minutes).