Observability Assessment Questionnaire

0/13

How healthy is your observability practice?

See where your strongest and weakest areas are, how you compare to enterprise patterns we’ve seen across
financial services, retail and federal environments, and what the highest-leverage next step would be. No prep
required. Most people answer from memory. If you have to guess on a question, that’s a signal worth noting.

  • 13 Questions
  • 5 Minutes
  • No Sales Pitch

Question 1 : Tool Landscape

How many distinct monitoring or observability tools are actively in use across your environment today?

  • 1 to 3 tools, consolidated, intentional choices
  • 4 to 6 tools, some overlap but mostly intentional
  • 7 to 10 tools, significant overlap, hard to map who uses what
  • 11 to 20 tools, accumulated through acquisitions and team preferences
  • More than 20 tools, or genuinely don't know

Question 2 : Cost Visibility

Could you, within an hour, tell me what your annual observability spend is across all platforms combined?

  • Yes, we track this monthly with chargeback to teams
  • Yes, we know the number but don't break it down by team or service
  • Roughly, I could estimate within 20%
  • No, I'd need to pull data from finance and multiple vendors
  • No, and I suspect the answer would surprise leadership

Question 3 : Cost Growth

How has your observability spend trended over the last three years?

  • Flat or declining, we've actively managed it
  • Grown roughly in line with infrastructure or business growth
  • Grown 50–100%, outpacing expected growth
  • Grown 100–300%, and we're not sure why
  • More than tripled, leadership has asked questions we haven't answered

Question 4 : Instrumentation Approach

What instrumentation approach are your applications using today?

  • OpenTelemetry across most or all services, vendor-agnostic by design
  • Mixed: OpenTelemetry adoption underway, with some vendor SDKs still in place
  • Mostly vendor-specific SDKs (Datadog, New Relic, Dynatrace, Splunk) with no OTel migration plan
  • Mostly vendor-specific SDKs and we know we have a lock-in problem
  • Inconsistent, some services well-instrumented, others not at all

Question 5 : Alert Quality

What percentage of alerts your team receives result in meaningful action?

  • More than 80%, we've actively managed alert quality
  • 50–80%, manageable but room to improve
  • 20–50%, we triage a lot of noise
  • Less than 20%, alert fatigue is a recognized issue
  • Honestly, I don't know, and that's part of the problem

Question 6 : Data Correlation

When a critical incident happens, how quickly can you correlate metrics, logs and traces for the affected service?

  • Under five minutes, our telemetry is correlated by trace ID across systems
  • 5 to 30 minutes, we know where to look but it takes a few clicks
  • 30 to 90 minutes, different tools, manual correlation, tribal knowledge required
  • Hours, we usually find the answer in postmortem rather than during the incident
  • We typically don't correlate at all; we resolve symptoms and move on

Question 7 : On-Call Health

How would your on-call engineers describe their current rotation?

  • Sustainable, most weeks have no major incidents and false pages are rare
  • Manageable, occasional bad weeks but generally healthy
  • Stressful, recurring noise, frequent paging, accepted as normal
  • Burning people out, engineers ask to leave the rotation
  • We don't have a formal rotation; whoever is around handles things

Question 8 : Dashboard Usage

What percentage of dashboards in your primary observability tool are actively used (viewed in the last 30 days)?

  • More than 70%, we audit and prune unused dashboards regularly
  • 40–70%, most dashboards are used but some clutter exists
  • 20–40%, significant dashboard sprawl from years of accumulation
  • Less than 20%, we have hundreds of dashboards and most are abandoned
  • I don't know how to even check this

Question 9 : Observability Ownership

Who owns observability strategy in your organization today?

  • A dedicated platform or SRE team with executive sponsorship
  • A platform team with shared responsibility, strategy is informally driven
  • Distributed, each team owns their own observability with minimal coordination
  • Unclear, multiple groups claim ownership, none have authority
  • Nobody owns it, observability happens through individual heroics

Question 10 : AI Readiness

How positioned is your observability data foundation for AI-driven operations (anomaly detection, automated correlation, intelligent alerting)?

  • Strong, clean instrumentation, structured data, ready for AI features
  • Workable, we'd need to clean some things up but the bones are there
  • Mixed, pockets of good data, pockets of mess
  • Weak, our telemetry quality would make AI features unreliable
  • We haven't thought about this and don't have a position yet

Question 11 : AI Systems in Production

Are you running LLM-powered applications or autonomous agents in production, and can you see what they're actually doing?

  • We run AI in production with instrumentation for token cost, latency, output quality and failure modes
  • We run AI in production with basic telemetry, but not answer quality or drift
  • Not running AI in production yet, though it's on the roadmap
  • We run AI in production and honestly can't see much beyond whether it's up
  • AI is being deployed by teams and we don't have visibility into where or how

Question 12 : Vendor Renewal Pressure

When does your largest observability vendor contract come up for renewal, and what's expected?

  • More than 12 months out, no immediate pressure
  • 6–12 months out, we're starting to evaluate options
  • Less than 6 months out, actively evaluating alternatives
  • In progress now, and the price increase is significant
  • Recently renewed at a steep increase and we feel locked in

Question 13 : Readiness To Act

If a clear roadmap and quantified savings opportunity were on your desk in the next 60 days, what would happen?

  • We'd move forward immediately, budget and authority are aligned
  • We'd take it to leadership with a strong likelihood of approval
  • We'd evaluate it seriously but it would compete with other priorities
  • It would need to wait until the next budget cycle
  • We're in a frozen state, no major initiatives until further notice

Results

0/52