Enterprise · Long-form Report
Why Observability Costs Grow Faster Than the Systems They Watch
Telemetry was supposed to make operations cheaper to understand. Instead, logs, metrics, and traces accumulate their own economics, and the bill rises even on weeks when traffic stays flat.

Independent coverage
Contributing Writer — AI / Data / Business · Freelance
Edited by Dr. Elin Lindqvist, MD
Published 29 September 2026
10 min read
Evidence: Analysis
Observability platforms promised a simple trade. Pay a modest amount for telemetry, and never again debug a production system blind. In most enterprises the trade did not hold. The telemetry line item now grows quarter after quarter, frequently faster than the infrastructure it describes, and engineering leaders find themselves negotiating with their own monitoring tools the way they once negotiated with cloud providers.
The pattern repeats across industries and company sizes. A platform team adopts structured logging, distributed tracing, and high-resolution metrics. Dashboards improve. Incident response improves. Then the invoice arrives, and the numbers do not match any intuitive model of what the systems do. Traffic grew ten percent. The observability bill grew sixty. The discrepancy is not a pricing error. It follows from how telemetry is produced, retained, and billed.
Telemetry is billed like a product, not a utility
Compute and storage are priced by capacity. Telemetry is priced by volume of events, gigabytes ingested, or hosts monitored, and each of those units scales with something the engineering team does not directly control. A verbose library upgrade, a new microservice, a marketing campaign, or a retry storm all change the ingest volume without any deliberate observability decision.
This makes observability one of the few infrastructure categories where the bill is a function of code behaviour rather than provisioned capacity. A team can reduce compute spend by shutting down instances. Reducing telemetry spend requires changing what the software emits, which means touching application code, configuring agents, and negotiating with every team that depends on the dashboards those signals feed. The cost is centralised. The cause is distributed. That asymmetry is the root of the growth problem.
Every new service multiplies the signal
Modern architectures generate telemetry multiplicatively. In a monolith, one application emits logs and metrics. In a service-oriented system, every request crosses several services, and each hop emits its own spans, logs, and metric series. Add a service, and you add not one emitter but a new set of cross-product interactions with every existing emitter.
Tracing makes this explicit. A single user request through a fifteen-service path produces hundreds of spans, each carrying attributes, timestamps, and identifiers. Multiply by request volume, and the trace pipeline processes more records than the business transactions it describes by several orders of magnitude. The engineering value is real, because those spans are often the only way to localise a latency problem. But the cost structure means the observability bill tracks architectural fan-out, not business growth.
High cardinality is the expensive part
Metrics are cheap in aggregate and expensive in dimension. A metric that records request latency with labels for endpoint, region, status code, customer tier, and version does not cost one time series. It costs one series per unique combination of label values. Teams that add a customer identifier or a request identifier as a label discover this the hard way: a single poorly chosen dimension can multiply the series count from thousands to millions.
Cardinality incidents are now a recognised operational failure mode. A developer adds a high-cardinality label to debug a problem, ships it, forgets it, and the metrics backend quietly absorbs the load for months. Some platforms throttle or drop excess series, which silently blinds the dashboards that depended on them. Others pass the cost through. Either way, the person who caused the growth rarely sees the consequence, because the bill lands somewhere else.
Retention defaults outlive their usefulness
Most observability data is read rarely and written expensively. Studies of operational behaviour, and the consistent experience of platform teams, show that the overwhelming majority of log lines are never queried after the first hours of their life. Yet default retention settings commonly keep full-fidelity logs for thirty, sixty, or ninety days, and keep the storage tiering configured at ingestion time rather than reviewed against actual query patterns.
Retention drift compounds quietly. A compliance requirement from one project sets a two-year retention on one index. A debugging exercise sets verbose logging on one service and nobody reverts it. An acquisition adds an entire estate with its own telemetry configuration. Each decision was locally reasonable. Together they produce a system that pays premium storage prices for data that answers questions nobody asks anymore.
Sampling is deferred because it feels like losing data
The standard remedy for trace volume is sampling: keep a representative fraction of traces, and keep all of the traces that matter, such as errors and slow requests. The technique is well understood and built into every major tracing standard. It is still underused, because sampling feels like a loss of completeness. Engineering culture treats telemetry as evidence, and discarding evidence feels irresponsible even when the discarded portion is redundant.
The deferral has a cost of its own. Teams that ingest everything pay full price for the privilege of never sampling, and then discover during a major incident that the volume they paid for slows down the very queries they need. Head-based sampling of routine traces, combined with tail-based retention of the interesting ones, delivers nearly all of the investigative value at a fraction of the ingest. Adopting it later means retrofitting instrumentation across a dozen teams, which is why the decision keeps getting postponed.
The bill lands on a different budget than the cause
Observability spend is usually paid centrally, by a platform or infrastructure group, while the decisions that generate the spend are made by application teams. The central team has the incentive to cut volume. The application teams have the incentive to keep their dashboards rich and their debugging easy. Without a mechanism that connects the two, the volume grows and the negotiation is adversarial.
Organisations that control this well do three things. They attribute telemetry cost to the teams that generate it, so the trade-off is visible to the people making it. They set default emission standards in shared libraries and agents, so the compliant path is also the cheap path. And they review telemetry configuration like any other production resource, with the same rigour applied to compute and storage. None of this requires new technology. It requires treating telemetry as a product with a cost of goods, rather than a background service.
What disciplined telemetry looks like
The operations that keep observability costs proportional share a few habits. They define, per signal class, how long full fidelity is actually needed, and tier storage to match. They enforce cardinality budgets in the instrumentation layer instead of discovering violations on the invoice. They sample traces by default and unsample deliberately. They delete dashboards nobody opens, because every panel is a standing query. And they periodically answer the hardest question in the discipline: which of this data changed a decision in the last quarter, and which of it merely exists.
The goal is not minimal telemetry. Under-instrumented systems cost more than over-instrumented ones, just in a currency of incident hours instead of invoices. The goal is deliberate telemetry: signals that are emitted because someone will read them, retained because someone will need them, and priced where someone can see the cost of emitting them. Observability earns its keep when the questions it answers are worth more than the pipeline costs to run. Keeping that equation honest is an engineering discipline, not a procurement decision.
"Every new service multiplies the signals it emits, and every signal is billed. The system grows linearly. The bill grows combinatorially."
Sources
Related reading