The Observability Contract Trap
By Brad Bell, Principal Observability Lead
Two years ago your team ran a proof of concept, the numbers looked great, and you signed a multi-year observability contract at a rate everyone was happy with. Now the renewal is in front of you, the bill has tripled, and nobody did anything wrong. That’s the trap, and it catches good teams precisely because the mistake isn’t visible at signing time.
The observability contract trap is signing a multi-year deal priced against pilot-scale telemetry, when the pricing model scales with your telemetry volume rather than the value you get from it. The rate you negotiated is almost irrelevant to what you’ll actually pay. The thing that determines your bill is the growth curve of your telemetry, and that curve is steeper than almost anyone models at signing. Cost has been the top tool-selection criterion in Grafana’s Observability Survey for three years running, cited by 65% of respondents in 2026, and this is a big part of why: teams keep getting surprised by the second number, not the first.
This is not a story about any vendor behaving badly. The major commercial platforms are powerful and their pricing is usually transparent about how it works. The trap is structural, it lives in the mismatch between how these tools are priced and how telemetry grows, and understanding that mismatch is how you avoid walking into it.
Why the bill outruns the deal
Observability platforms generally price on some measure of volume: hosts, ingested gigabytes, custom metrics or active series, spans, or some combination. Every one of those measures grows as your system does. Ship more services, and hosts and spans climb. Add more dimensions to your metrics, and series count climbs. Turn up log verbosity to debug an incident and forget to turn it back down, and ingest climbs. The pricing scales with the telemetry, and the telemetry scales with your success.
A proof of concept runs against a slice of production, for a few weeks, with whatever instrumentation you had at the time. It is, almost by definition, the smallest and cleanest your telemetry will ever be. The moment you roll the platform out for real, three things start pushing the volume up at once, and none of them show up in the POC.
The three multipliers nobody models
The first multiplier is footprint growth. More services, more hosts, more containers, more environments. This one teams at least partly anticipate, though they usually underestimate it, because a healthy engineering organization ships, and every new thing it ships emits telemetry.
The second multiplier is cardinality, and it’s the quiet killer. Every new label, every high-uniqueness dimension like a user ID or a request ID attached to a metric, multiplies the number of time series you’re paying to store and query. Cardinality doesn’t grow linearly with your system, it can explode from a single well-intentioned code change, and on per-series pricing that explosion lands directly on the bill.
The third multiplier is verbosity and retention creep. Debug logging left on after an incident. Retention windows set generously “just in case” and never revisited. Sampling that was supposed to be tuned and never was. Each of these quietly raises the volume floor, and because they accumulate gradually, no single change ever looks like the problem.
Stack these three on top of a rate negotiated against pilot volume, and you get the renewal surprise. The deal was fine. The growth was the story.
Why negotiating harder doesn’t save you
The instinct at renewal is to push for a bigger discount, and it’s the wrong lever. Squeezing another ten or fifteen percent off the rate does nothing about the volume that tripled underneath it, and next year the same curve keeps climbing off a slightly lower base. You optimized the price per unit while the number of units ran away from you.
This is the core insight that changes how you should approach the contract: your observability bill is downstream of your consumption, and consumption is an engineering and governance problem, not a procurement one. A better rate on unmanaged growth is a slower way to arrive at the same place.
What to do before you sign
Model the growth, not the snapshot. Take your POC volume and project it against your actual roadmap, the services you plan to ship, the environments you plan to add, the cardinality your instrumentation will realistically produce. Sign against where you’ll be in eighteen months, not where you are during the pilot, and you remove most of the surprise.
Get consumption controls in the contract and in your architecture. Before you commit, know how you’ll see per-team and per-service consumption, how you’ll cap or alert on runaway growth, and where you’ll filter and sample. A collection layer built on an open standard like OpenTelemetry lets you decide what’s worth keeping before it hits the metered backend, which is the single most durable cost control you can put in place.
Negotiate on the terms that actually bite. The unit rate matters less than the ramp, the overage treatment, and your ability to true up rather than get penalized as you grow. A contract that assumes healthy growth and prices for it fairly is worth more than a low headline rate with a punishing overage curve.
What to do if you’re already in it
If you’re reading this at renewal with the surprise already on the table, the move is the same one that would have prevented it: optimize first, then decide. Get visibility into what you’re actually collecting and who owns it. Attack the cardinality that’s inflating your series count. Rationalize retention down to what anyone actually queries. Move filtering and sampling to the collection layer. Teams that do this work routinely take meaningful volume, and therefore meaningful cost, back out of the platform, and the reductions are real even though the exact figure depends on your own consumption profile.
Only after that should you decide whether the platform itself is the problem. Often it isn’t, and the optimized bill is one you can live with. Sometimes it is, and now you have a data-backed case for a change instead of a reaction to a scary invoice. Either way, you’re deciding from the consumption side, which is the only side that actually moves the number.
The contract trap isn’t really about contracts. It’s about signing up for a volume-priced service without a plan to govern the volume. Put the governance in place, and the contract takes care of itself.
If you want to model your real telemetry growth before a renewal, or pull volume back out of a platform you’re already committed to, that’s exactly the kind of consumption analysis our assessment is built to run. It starts with your data, not a discount.
Learn more here!
About the Author
Brad Bell is the Principal Observability Lead at TekStream Solutions, a Digital Resilience partner that helps organizations operate, recover, and adapt with confidence by modernizing, securing, and optimizing their digital environments. His 29 years of experience spans large-scale telecommunications infrastructure, enterprise cloud strategy, and operational transformation across industries ranging from financial services to quick-service restaurants — work that included scaling monitoring platforms to 12,000+ nodes, architecting next-generation networks, and turning reactive ops teams into reliability engineering organizations. A Grafana Certified Solutions Architect and PreSales Solution Architect, he teaches SRE immersion courses at major financial institutions and consults on observability strategy at the enterprise level, helping clients navigate the convergence of cost pressure, tool sprawl, and AI-driven operations that is reshaping how enterprises think about reliability. His perspective isn’t theoretical — it comes from having been the person on the other end of a 2 AM page, and from watching organizations make expensive, avoidable mistakes with their observability investments. At TekStream, his focus is helping enterprises build platforms that solve today’s problems and hold up against where the industry is headed.