eBPF Observability with Beyla: Instrument Everything Without Touching the Code

By Brad Bell, Principal Observability Lead

You want traces and golden-signal metrics from your two hundred services. The problem is the instrumentation backlog: a third of those services are older code nobody wants to recompile, another third are in languages your team hasn’t gotten to yet, and manually instrumenting the rest is a multi-quarter slog that competes with actual feature work. So most of the fleet stays dark, and you find out about problems the slow way.

eBPF observability solves the first, hardest step of that problem. It gives you zero-code instrumentation, request-rate, error, and duration metrics plus basic traces, from services in effectively any language, without changing a line of application code or restarting anything. Grafana’s Beyla is the on-ramp, and as of 2025 it’s also the foundation of a broader open standard: Grafana donated Beyla to the OpenTelemetry project, where it now anchors OpenTelemetry eBPF Instrumentation, known as OBI. Beyla continues as Grafana’s distribution of that upstream project. eBPF is not a silver bullet, and the honest answer for a serious estate is a hybrid one, but as the fastest way to go from blind to a baseline, it’s hard to beat.

eBPF lets you run small, safe programs inside the Linux kernel, attached to specific events, without modifying the kernel or the applications running on it. Applied to observability, that means a tool can watch the network and system calls an application makes and reconstruct what it’s doing, from the outside, without the application knowing or caring.

Beyla, and OBI underneath it, use this to capture RED metrics, request rate, errors, and duration, and basic trace spans for the traffic a service handles. Because the instrumentation happens at the protocol level in the kernel rather than inside a language runtime, it works across languages: Go, Java, .NET, Python, Ruby, Node.js, Rust, C and C++, and more, all from the same tool. There are no code changes, no redeploys, no new application dependencies, and therefore no new dependency-related security surface inside your app. It runs on any reasonably modern Linux, kernel 5.8 or later, and in Kubernetes it typically runs as a DaemonSet so one instance per node covers everything on that node.

The practical effect is that you can drop it into a running cluster and have a service-level health baseline, and a service graph showing who calls whom, for your whole fleet in an afternoon, including the services nobody was ever going to hand-instrument.

This is worth understanding, because it speaks to how you should think about betting on it. Grafana built Beyla starting in 2023, and it grew into one of the more capable eBPF instrumentation tools available. Rather than keep it proprietary, Grafana donated it to OpenTelemetry in 2025, where it became OpenTelemetry eBPF Instrumentation, developed in the open with contributors from Grafana, Splunk, Coralogix, Odigos, and others. Development has accelerated under the shared project, and it reached its early releases through 2026 on the road to a 1.0.

For an enterprise making a long-term bet, this matters. The core technology is now a vendor-neutral, community-owned OpenTelemetry project, not a single vendor’s lock-in play. Beyla remains Grafana’s distribution of it, adding conveniences like tighter Grafana Cloud onboarding and integration with Grafana Alloy, but the foundation is open. You get the Grafana on-ramp without betting your instrumentation strategy on one company’s roadmap, which is exactly the kind of position we like to see a client in.

eBPF instrumentation earns its place in a few specific situations. It’s the fastest possible way to establish a baseline across a large or heterogeneous fleet, because it doesn’t care what language anything is written in. It’s the answer for legacy and compiled services that can’t or won’t be recompiled to add an SDK. It’s excellent at the things that live at the network and process level, service graphs, network metrics between services and availability zones, and process-level resource usage, which SDKs don’t give you out of the box. And it produces consistent telemetry across every service it touches, because the same tool is generating it the same way everywhere, instead of each team’s instrumentation drifting in its own direction.

Adoption reflects this. The CNCF’s observability group reported in early 2026 that a large majority of production Kubernetes clusters are already running at least one eBPF-based observability tool, which is a striking jump from just a couple of years earlier when eBPF was mostly a thing teams were talking about rather than running.

Because this is a mastery piece and not a pitch, here’s where eBPF stops.

Distributed tracing is genuinely hard at the eBPF level. Capturing that a request happened is straightforward, but stitching requests into full distributed traces requires propagating trace context, which means touching request payloads, and that’s difficult to do from the kernel, especially with encrypted traffic. The tooling applies clever heuristics and keeps improving, but deep, reliable distributed tracing is still where language SDKs lead. eBPF also doesn’t give you logs, and it doesn’t capture the custom, business-level spans and attributes that make a trace actually useful for debugging your specific application, the “which tenant, which feature flag, which customer tier” detail that only your code knows to emit. And there’s an operational caveat worth knowing up front: on clusters already running eBPF-heavy networking like Cilium or Calico’s eBPF dataplane, instrumentation probes can collide with the network layer’s programs, and the recommended pattern is to split responsibilities rather than have both fight over the same ground.

The mistake is treating eBPF and SDKs as competitors and picking one. They’re complementary, and the strong pattern, the one Grafana itself recommends, is to run both.

Use eBPF to instrument the whole fleet for breadth: an immediate baseline, service graphs, network and process metrics, and coverage of everything you were never going to hand-instrument. Use SDKs on the services that need depth: the critical paths where you want reliable distributed tracing and rich, custom, business-level context. The tooling is built for this coexistence, OBI and Beyla detect when an SDK is already reporting and back off to avoid duplicating signals, so you don’t pay twice or double-count. You end up with wall-to-wall baseline coverage from eBPF and deep instrumentation exactly where it earns its cost, which is a far better place than the all-or-nothing instrumentation backlog you started with.

In a Grafana stack, the flow is clean: Beyla or OBI emits OpenTelemetry and Prometheus data, ships it through Alloy over OTLP, and lands it in Tempo for traces and Mimir for metrics, where it correlates with everything else in Grafana. Zero-code breadth feeding the same backends as your deliberate, SDK-based depth.

Start with eBPF to stop flying blind across the fleet. Layer SDKs where the depth pays for itself. That sequence gets you visibility in an afternoon and quality over time, which is the right order to buy them in.

Standing up eBPF-based baseline coverage and then layering SDK depth where it matters is concrete, high-leverage work, and it’s the kind of thing our practice builds alongside your team. If half your fleet is dark, that’s a good first place to look.

Learn more here!

About the Author

Brad Bell is the Principal Observability Lead at TekStream Solutions, a Digital Resilience partner that helps organizations operate, recover, and adapt with confidence by modernizing, securing, and optimizing their digital environments. His 29 years of experience spans large-scale telecommunications infrastructure, enterprise cloud strategy, and operational transformation across industries ranging from financial services to quick-service restaurants — work that included scaling monitoring platforms to 12,000+ nodes, architecting next-generation networks, and turning reactive ops teams into reliability engineering organizations. A Grafana Certified Solutions Architect and PreSales Solution Architect, he teaches SRE immersion courses at major financial institutions and consults on observability strategy at the enterprise level, helping clients navigate the convergence of cost pressure, tool sprawl, and AI-driven operations that is reshaping how enterprises think about reliability. His perspective isn’t theoretical — it comes from having been the person on the other end of a 2 AM page, and from watching organizations make expensive, avoidable mistakes with their observability investments. At TekStream, his focus is helping enterprises build platforms that solve today’s problems and hold up against where the industry is headed.