Add stack and space data to drift section of GraphQL API
Hi, our team has created a PR (running locally in our LGTM stacks) of the promex service that adds space and stack labels for metrics that have that data available. We would like to see additional data for stack and space added to drift data in the GraphQL API so that we have this available for our dashboards etc.
- Problem
Log in to comment and vote
Comments7
Krzysztof Niepokojczycki
Nov 24, 2025
Hey @Matt Becker let me chime in. I’m trying to understand the problem in more depth. As I see it, the current request will lead to a cardinality explosion, as stack (and other entities) IDs are unique. Implementing it might lead to a vast increase in the processing cost of the data series.
Could you elaborate on what you are trying to achieve? Maybe we can propose some sort of workaround (or see how we can’t and what we can do to unlock it) tuned in specifically for your needs, but I don’t see us currently moving forward with adding that to the Prometheus exporter.
Apricot Koala
Nov 24, 2025
Hi Krzysztof,
Thanks for reaching out. There will be an increase in cardinality, but a useful one. Currently the metrics for failed or drifted stacks don’t give us any detail on which stacks fail or drift, which isn’t terribly useful. We have dashboards that utilize these metrics and being able to see the stack and space detail historically would help us greatly in tracking stack issues for our teams.
I’ve got a PR that adds support for this to the promex service already, but need some additional GraphQL API data as outlined below in the queries I sent to Marcin. The queries below also include support for stacks awaiting approval, another metric we would like to observe.
Regards,
Matt
Krzysztof Niepokojczycki
Dec 4, 2025
Hi @Matt Becker,
Unfortunately, we can’t increase the cardinality with these dimensions due to the massive increase in maintenance cost, especially for customers that operate on huge amounts of stacks or use ephemeral stacks. The metrics’ current objective is to provide insight into the number of stacks/resources that are generally in a given state, rather than identifying singular stacks that have drifted.
However, I recognize your need to somehow rely on the drifted stacks information in your LGTM setup. We can propose creating a separate metric for it, so that other consumers of existing metrics remain unaffected. Would that work for you?
Black Breeze
Nov 20, 2025
In an ideal world the exporter would not need to go through the GraphQL API, which is one giant n+1 factory. At some point we’d like to have dedicated endpoint for exportable metrics that are optimized for frequent reads. But for now as a stopgap solution I’m happy to have the team investigate performance implications of adding extra GraphQL resolvers.
Can you please provide the exact GraphQL queries you’d like to be able to make?