Richer native Prometheus metrics: full exporter coverage, run timing, worker/queue metrics
It would be really nice to get metrics on # runs, run success, # stacks, etc. via Prometheus so we can create custom internal dashboards with this info. It’s helpful for measuring adoption, but we don’t use Datadog so we can’t get the info from there.
- Workaround
Log in to comment and vote
Comments22
Natalia Gazda
Dec 13, 2024
Input from @Joost Polley
“there's a lot of things we'd like to see actually:
more in depth metrics about stacks (run duration & such)
amount of triggered runs, and then data on the end state of those runs
and much more I can't think of at the moment
I do know about the dashboard feature and it's great, but we're all nerds and we really appreciate the data in prometheus / grafana next to all our other infrastructure metrics ...So I hope this feature does not disappear, but actually gets more features”
Blue Goose
Dec 13, 2024
Another useful metric would be the amount of stacks that are using a certain vendor (Terraform, OpenTofu, CloudFormation, Ansible, …).
Amaranth Dew
Dec 13, 2024
To expand on the above, with Atlantis (which we’re moving away from), the prometheus metrics include things like Terraform version, repository, success/failure, time spent on init vs. plan vs. apply phases - all of which lets us run various analyses to understand how we use Atlantis (and is how we were able to analyze our Atlantis usage to map to # of Spacelift workers needed)
Jonah Kowall
Jul 12
We’ve merged several closely related requests into this one so votes and discussion live in one place. The consolidated ask, as we understand it: a first-class native Prometheus/OpenMetrics endpoint (or exporter v2) with coverage at least at parity with the Datadog integration, documented, and efficient enough to scrape frequently. Specific metric families requested across the merged posts:
Run lifecycle: end-to-end run duration histograms (p50/p90/p99) with labels for stack, space, run type, terminal state; runs by state/type; unconfirmed/retryable runs; drift detection runs.
Worker pools: busy/idle workers, queue length (gauge), per-run worker wait time (histogram), utilization over time.
Stacks and governance: stack states (failed, drifted, locked, disabled), vendor and tool-version usage across stacks.
Exporter mechanics: full coverage of the metrics API, adjustable histogram buckets, docs for every metric.
This maps to the observability improvements we have referenced on our roadmap. Keep voting and add missing metrics in the comments; specifics here directly shape scope.
Jonah Kowall
Sep 22
This is a really demanded feature that we've been tracking for quite some time. I've got good news to share with you all.
v0.12.0 is a prerelease build. We are looking for feedback before making it official, so if you hit a bug, a regression, or a metric that looks wrong, please open an issue at https://github.com/spacelift-io/prometheus-exporter/issues.
If you monitor Spacelift with Prometheus, this release makes the exporter faster, more transparent about what it is doing, and easier to tune to your account.
Missing metrics are back.
spacelift_worker_pool_workersandspacelift_scrape_duration_secondswere defined but never exposed. Both now show up on the/metricsendpoint.Scrapes run concurrently. Collectors used to query the Spacelift API one after another, so a slow endpoint held up everything. They now run in parallel, which cuts total scrape time and lowers the risk of Prometheus timeouts on larger accounts.
Per-collector health. Each collector now reports its own success status and request duration. When a scrape degrades, you can see which collector is slow or failing instead of guessing from a single aggregate number.
Partial scrapes, opt in. Previously, one failing collector failed the whole scrape and you lost every metric. You can now opt into partial scrapes so healthy collectors keep reporting while the broken one is flagged. Collectors that hit features your account does not have are marked unsupported rather than failed, so they stop generating noise.
Turn collectors on and off. New
--collector.<name>and--no-collector.<name>flags let you disable collectors you do not need. Fewer API calls, faster scrapes, and no metrics you never look at.Under the hood. Each collector now issues a single named GraphQL operation, which makes API-side tracing and debugging easier. The project also gained a collector test harness with golden metric snapshots, so metric names and labels are locked down against accidental changes. Dependencies were bumped, including prometheus/client_golang 1.24.1 and urfave/cli v3.12.0.
Thanks to first-time contributors @aerfio and @jkowall.
Full release notes and binaries: https://github.com/spacelift-io/prometheus-exporter/releases/tag/v0.12.0
Sapphire Chipmunk
Jan 9
I’d really appreciate if this topic could be revisited and given higher priority.
While exploring the GraphQL API, I confirmed there is already enough high-value data available today to expose much richer Prometheus metrics without guesswork. I documented a realistic set of candidate metrics covering:
Vendor & tool usage (Terraform vs OpenTofu, Terragrunt usage and versions)
Stack state & governance signals (disabled, blocked, locked, protected from deletion, drifted)
Run lifecycle signals (runs by state/type, unconfirmed runs, retryable runs, drift detection runs, run lag)
Worker pool signals (users, limits, management model)
These metrics would directly address the gaps mentioned in this thread and unlock better capacity planning, adoption tracking, and operational visibility for teams using Prometheus and Grafana.
Given that the API already exposes the underlying data, I believe prioritizing this would deliver significant value to many customers.
Green Rice
Jul 12
•Merged request
•1 vote
Export Additional Metrics to spacelift-promex
We are heavy users of Grafana and Prometheus for observability and alerting. The current metrics that Spacelift exposes to prometheus today is very limiting and we want to build out a rich set of dashboards that give insight into platform performance and also the developer experience.
We noticed that only ~5/36 graphql metrics are being exposed in the spacelift-promex
Asks:
Ensure that spacelift-promex is exporting all metrics available in the metrics API
Be able to adjust histogram buckets for all relevant metrics
Salmon Potato
Jul 12
•Merged request
•6 votes
Spacelift Prometheus Exporter Being Actively Developed?
Are there any plans to add metrics / pivots around spaces and individual stacks?
Black Breeze
Jun 8, 2025
Natalia Gazda
Jun 9, 2025
Hey @Julio Chana, thanks for the requests! I’m currently looking at the observability topic and would love to understand your use cases a bit more.
What decisions or alerts are you trying to drive with these metrics?
Who uses the metrics now, and who would use the new ones?
If you had these metrics, what would become easier or faster?
White Toad
Jul 12
•Merged request
•6 votes
Export Spacelift operational metrics to Prometheus / Grafana, richer metrics beyond current exporter
We use Grafana and Prometheus heavily for day to day monitoring and alerting. Today Spacelift has dashboards in the UI and there is a Prometheus exporter, but the metric coverage is limited and the UI is mostly something we open only when something is already broken.
We want a standard way to export richer Spacelift operational metrics so we can proactively detect regressions and performance changes, for example when a workflow suddenly becomes slower over time, or when a specific stack starts taking longer after a change.
A concrete example. We recently spent time optimizing Ansible execution time. We want to make sure it does not regress in the future. Having metrics we can store and alert on would help a lot.
What we want (ideal solution): A Prometheus compatible endpoint or exporter with broader, well documented metrics, so we can scrape them and build Grafana dashboards and alerts. Alternatively, a direct Grafana integration is fine, but a standard metrics endpoint is best.
Black Sorbet
Jul 12
•Merged request
•12 votes
Expose Busy / Queue length Worker Pool Metrics
I would like to be able to systematically monitor our worker pool to get quantitative data on our internal developer experience as our engineers all share the worker pool to run their IaC.
I would like to use those metrics to improve decision making on the number of workers we need for our organizations needs.
Black Breeze
Mar 4
Hey Loic, thanks for raising this. We have a roadmap item for later this year to look into a more comprehensive overhaul of our observability story. I will flag this accordingly and open up for community voting to gauge overall interest in these metrics once we dive deeper into this particular topic.
In the meantime, I think a good approximation would be tracking the time that runs spend in the READY state, which is something a notification policy can help you with. Not sure what vendor you’re using but our Datadog integration is a good start, even if to understand the logic behind the metric calculations.
Plum Pen
Apr 7
We would love to see this as well - especially if it was integrated into the datadog integration ya’ll provide.
pending runs and queueed runs - each tagged per workerpool? or something along those lines.