When You Should NOT Leave Your Commercial Platform

By Brad Bell, Principal Observability Lead

Your renewal is six months out, the observability line item has doubled since the last one, and someone in a leadership meeting has already said the words “why don’t we just move to something cheaper.” Here is the advice almost nobody in a sales cycle will give you: sometimes the right move is to stay.

The decision to leave a commercial observability platform should follow the data, not the frustration. For a meaningful share of the teams we talk to, the honest answer after the analysis is that optimizing the platform they already own beats the cost, risk, and disruption of migrating off it. Optimize first, then decide. That order matters, because a migration you start from a place of frustration tends to carry the same problems into the new tool and hand you a fresh bill on top.

This is not a defense of any vendor. Datadog, Splunk, Dynatrace, and New Relic all do genuine work well, and all of them can produce a bill that outruns the value if nobody is watching the consumption side. The point is narrower and more useful: switching platforms is one of the most expensive things an engineering organization can do, and you should be sure the pain you feel is actually a platform problem before you pay that price.

Is the cost problem the platform, or how you’re using it?

This is the first question, and it decides most of the others. When an observability bill climbs faster than the business, the cause is usually one of three things, and none of them is the vendor’s logo.

The first is consumption patterns nobody anticipated when the contract was signed. A team ships more services, cardinality quietly explodes, a debug logger gets left on in production, and the ingest curve bends upward while nobody is looking. The second is a lack of governance, meaning no one owns the question of what telemetry is being collected, why, and whether anyone has looked at it since. The third is an architecture decision that routes everything to an expensive backend without filtering, sampling, or tiering along the way.

Here is why this matters for the stay-or-go decision. If your pain traces back to consumption, governance, or architecture, then migrating platforms does not fix it. It relocates it. The same uncontrolled cardinality that made your current bill hurt will make your next one hurt, and now you have paid a migration tax for the privilege. Grafana’s 2026 Observability Survey, which drew on 1,363 responses across 76 countries, found that cost has been the single most important tool-selection criterion for three years running, cited by 65% of respondents. That tells you the whole market is feeling this. It does not tell you that switching is the answer, because the survey also shows the pain concentrates around complexity and signal-to-noise, which travel with you.

So before anything else, get honest about whether you have ever actually optimized the platform you have. Most teams have not. They have paid the bill, complained about the bill, and started shopping. The optimization step, the one that comes before the decision, is the part that usually gets skipped.

What does the switch actually cost?

The sticker comparison, your current annual spend against a cheaper platform’s list price, is the most misleading number in this entire decision. It ignores the cost of the move itself, which is where migrations quietly go over budget.

A real migration cost model accounts for several things at once. There is the engineering effort to re-instrument services, rebuild dashboards, and rewrite alerts, which for a large estate is measured in engineer-months, not weekends. There is the period of parallel running, where you pay for both platforms at the same time because you cannot cut over cold, and that overlap often lasts longer than planned. There is retraining, because your on-call engineers are fluent in the tool they have and slow in the tool they do not. There is the productivity dip during the transition, when incident response is worse precisely because the muscle memory is gone. And there is the risk that something breaks during the move, in the one system whose entire job is to tell you when things break.

Add those up and compare them against the actual, modeled savings, not the list-price fantasy. Sometimes the math is clearly favorable and you should move. Often it shows a breakeven point two or three years out, which means the savings are real but the payback is slow, and slow payback changes how you sequence the decision. The number you want is the three-year total cost of both paths, including the migration, not the monthly delta between two price sheets.

What would you actually lose?

Name the strengths before you walk away from them. This is not sentiment, it is due diligence, because the strengths are what your team quietly depends on.

A well-run Datadog environment gives you a genuinely integrated experience across signals, and teams that value having everything in one place are not wrong to value it. Splunk’s search language and security analytics are hard to match, and if your compliance and detection workflows are built on them, that is load-bearing infrastructure, not a nice-to-have. Dynatrace’s automated dependency mapping does work that would otherwise fall to a human. New Relic’s pricing model fits some usage shapes better than others. Whatever platform you are on, there is a reason you are on it, and the version of you that is angry about the bill has usually forgotten what that reason was.

The question is not whether the platform is good. It is whether the specific strengths you rely on are ones you would have to rebuild, badly and slowly, somewhere else. If they are, that raises the bar the savings have to clear.

When staying and optimizing is the right call

Stay when the pain is fixable in place. If your bill is being driven by unused custom metrics, over-retained logs nobody queries, high-cardinality labels that snuck in, or duplicate collection across teams, those are optimization problems, and you can solve them without a migration. We regularly see organizations take meaningful cost out of a commercial platform purely by fixing what they send to it, and the reductions are real even though we frame them as an addressable range rather than a guarantee, because the actual figure depends on your own consumption profile.

Stay when the switching cost exceeds the savings over your planning horizon. If the three-year model shows the migration eating most of the benefit, the disruption is not worth it, and your energy is better spent on governance.

Stay when the platform’s strengths are load-bearing. If your security team lives in Splunk or your incident response is built around a Datadog workflow that works, the burden of proof for ripping that out is high, and “the bill is annoying” does not meet it.

When leaving is the right call

Honesty runs both ways, so here is the other side. You should move when the pain is structural rather than behavioral, meaning you have genuinely optimized and the economics still do not work for where you are headed. You should move when the pricing model punishes the exact thing your business does, so that success itself makes the bill worse in a way no amount of governance can offset. You should move when proprietary instrumentation has become the lever the incumbent uses to keep you in place, and breaking that lock is worth the cost of the move. And you should move when your future direction, more open standards, more portability, a different architecture, is one the current platform cannot follow you into.

The difference between this and the frustration-driven exit is that you arrive here after the analysis, not before it. You have optimized, you have modeled the switch honestly, you have named what you would lose, and the data still points to the door. That is a good migration. The one that starts with an angry renewal meeting usually is not.

How to optimize before you decide

If you take the “optimize first” step seriously, here is where the leverage is. Start with governance, because you cannot fix what you cannot see: get visibility into what telemetry you collect, who owns it, and whether anyone uses it. Attack cardinality, which is the quiet driver of metrics cost, by finding the high-cardinality labels that are inflating your series count without adding insight. Rationalize retention, because most organizations keep far more data, for far longer, than anyone ever queries. Move filtering and sampling to the collection layer, ideally on an open standard like OpenTelemetry, so you decide what is worth keeping before it hits the expensive backend rather than after. And put a cost owner in place, so the next surprise gets caught in a dashboard instead of a renewal.

Do that work, and one of two things happens. Either the bill comes down enough that staying is clearly right, or it does not, and now you have a genuine, data-backed case for moving that will survive contact with your CFO. Both outcomes are wins, because both replace a guess with a decision.

The teams that get this wrong treat the platform as the problem and the migration as the cure. The teams that get it right treat their own consumption as the problem, fix what they can, and let the numbers decide the rest. Optimize first, then decide.

If you want a structured read on whether your platform pain is fixable in place or genuinely structural, that is exactly what our assessment is built to surface. It starts with your consumption data, not a recommendation we picked in advance.

About the Author

Brad Bell is the Principal Observability Lead at TekStream Solutions, a Digital Resilience partner that helps organizations operate, recover, and adapt with confidence by modernizing, securing, and optimizing their digital environments. His 29 years of experience spans large-scale telecommunications infrastructure, enterprise cloud strategy, and operational transformation across industries ranging from financial services to quick-service restaurants — work that included scaling monitoring platforms to 12,000+ nodes, architecting next-generation networks, and turning reactive ops teams into reliability engineering organizations. A Grafana Certified Solutions Architect and PreSales Solution Architect, he teaches SRE immersion courses at major financial institutions and consults on observability strategy at the enterprise level, helping clients navigate the convergence of cost pressure, tool sprawl, and AI-driven operations that is reshaping how enterprises think about reliability. His perspective isn’t theoretical — it comes from having been the person on the other end of a 2 AM page, and from watching organizations make expensive, avoidable mistakes with their observability investments. At TekStream, his focus is helping enterprises build platforms that solve today’s problems and hold up against where the industry is headed.