Top open source APM tools (2026)
In open source application performance monitoring (APM), an OpenTelemetry Collector and a Helm chart can feed a columnar database. That setup provides distributed traces and correlated logs alongside request-level metrics with no license fee.
Self-hosting instead creates an operational burden because your team owns the dashboards and every part of the storage cluster, including upgrades.
This guide compares 11 open source APM tools on license, telemetry signals, deployment paths, and OpenTelemetry support, then walks through how to match a tool to your existing stack and the point at which self-hosting stops paying off.
What open source APM covers
APM tracks latency, error rate, and throughput across services and rolls those signals up into RED metrics and service level objectives.
OpenTelemetry’s primer defines a distributed trace as a record of the path a single request takes as it propagates through multiple services, and full APM renders those spans as service maps and flame graphs. The Cloud Native Computing Foundation (CNCF) draws the line at the long tail. Monitoring covers what hits most users, and observability lets you ask questions about the request nobody anticipated.
Open source APM tools fall into three architectural categories, and each one shapes how you deploy, query, and scale telemetry.
- Distributed tracing backends: Jaeger, Tempo, and Zipkin focus on trace data alone. They store and search spans and traces without covering metrics or logs.
- Metrics-first stacks: Prometheus and Grafana begin with time series and add other signals over time. Logs and traces layer in as separate components on top of the metrics foundation.
- Unified backends: SigNoz, Uptrace, Elastic APM, SkyWalking, and Pinpoint ingest logs, metrics, and traces through a single OTLP endpoint. Telemetry arrives over the OpenTelemetry Protocol (OTLP) into one system.
Which category fits comes down to what you already run in production.
The 11 best open source APM tools
Verify licenses in each project’s repository and hosted pricing on its pricing page before committing.
| Tool | License | Telemetry signals | Deployment | OpenTelemetry support | Best for |
| SigNoz | MIT Expat, except under SigNoz Enterprise License | Logs, metrics, traces | Docker, Helm, customer cloud | OTLP signal support | All signals, one store |
| Prometheus + Grafana (Loki, Tempo, Mimir) | Apache 2.0 (Prometheus) / AGPLv3 (Grafana, Loki, Tempo, Mimir) | Metrics, logs, traces (separate) | kube-prometheus-stack, Grafana Cloud | Tempo OTLP support | Prometheus teams adding signals |
| Jaeger | Apache 2.0 | Traces | OpenTelemetry Operator; beta Helm chart | Native OTLP support | Adaptive sampling on Kubernetes |
| Grafana Tempo | AGPLv3 with Apache-2.0 exceptions | Traces (Rate, Errors, and Duration metrics derived) | Object storage, Helm, cloud | Supported trace protocols | Object-storage trace retention |
| Elastic APM | Elastic License (binaries); AGPLv3 addition excludes APM Server (Elastic licensing details) | Traces, metrics, logs | Elastic Cloud on Kubernetes operator, self-managed | OTLP and EDOT intake | Existing Elastic Stack shops |
| Zipkin | Apache 2.0 | Traces | Single binary; no Helm chart | Zipkin OTLP bridge | Single-process tracing setup |
| Apache SkyWalking | Apache 2.0 | Traces, metrics, logs | skywalking-helm chart, self-hosted | OTLP receiver support | Java zero-code agents |
| Uptrace | AGPLv3 license | Traces, metrics, logs | Docker Compose, Helm, cloud | OTLP signal support | OpenTelemetry backend |
| Pinpoint | Apache 2.0 | Traces, Java Virtual Machine metrics | HBase; pinpoint-kubernetes chart | OTLP metrics support | Java method-level call trees |
| OpenObserve | AGPLv3 license | Logs, metrics, traces, real user monitoring | Single binary, Helm, cloud | OTLP signal support | Object-storage logs plus APM |
| Coroot | Apache 2.0 | Metrics, traces, logs, extended Berkeley Packet Filter profiling | Docker, Helm, operator, DaemonSet | OTLP and eBPF support | Kubernetes coverage without application instrumentation |
1. SigNoz
SigNoz is an OpenTelemetry-native backend for teams wanting logs, metrics, and traces in one user interface (UI). Its collector routes all three into ClickHouse storage, and SigNoz also manages a customer-cloud deployment in the customer’s account. Self-hosted teams operate the ClickHouse layer, while paid plans add managed and access-control options.
Pros
- Single query path: All three signals enter one UI.
- Vendor-neutral: OpenTelemetry ingestion avoids proprietary agents.
- Columnar storage: The architecture uses ClickHouse storage; inverted indexes are absent from this design.
Cons
- ClickHouse operations: Self-hosted teams manage the ClickHouse storage, including upgrades and backups as it scales.
- Paid access controls: SigNoz pricing places single sign-on in paid tiers.
- Deployment responsibility: The self-hosted architecture leaves database operations with the customer.
Pricing
The MIT core is free to self-host; Cloud Teams starts at $49 per month (promotional; $199 regular). Self-hosted teams remain responsible for ClickHouse infrastructure and operations.
Who is SigNoz best for?
Platform teams prepared to own a ClickHouse-backed observability service rather than operate separate signal stores.
2. Prometheus + Grafana (LGTM)
Prometheus + Grafana is a metrics-first stack for teams already scraping time series. Loki adds logs, Tempo traces, and Mimir long-term metrics behind one Grafana front end. Each signal retains its own storage and query components within the assembled stack.
Pros
- Established metrics model: Prometheus time series remain the starting point.
- Dashboard library: Dashboards import by ID.
- Incremental adoption: The LGTM components add signals to working metrics.
Cons
- Three query dialects: The stack uses three query languages across investigations.
- PromQL learning curve: Custom panels need hand-written queries.
- License review: The 2021 AGPLv3 relicensing can trigger legal review.
Pricing
Prometheus is Apache 2.0, Grafana and Loki licensing uses AGPLv3, and these tools are free to self-host. Grafana Cloud pricing describes the managed path.
Who is Prometheus + Grafana best for?
Prometheus teams willing to operate separate log and trace components instead of replacing their metrics foundation.
3. Jaeger
Jaeger is a tracing backend for Kubernetes teams that need sampling control. The project graduated from CNCF in 2019, and version 2 runs on the OpenTelemetry Collector framework. As a result, one binary can be the collector or ingester and can also be the query service.
Pros
- Sampling control: Sampling controls include adaptive and tail-based sampling.
- Storage choice: Storage documentation covers Cassandra, Elasticsearch, OpenSearch, and experimental ClickHouse.
- Collector reuse: Version 2 architecture is a Collector distribution.
Cons
- Traces only: The Jaeger architecture focuses on distributed tracing, so metrics and logs live elsewhere.
- Storage operations: Teams run the storage backend behind Jaeger.
- Version 1 finished: Jaeger v1 reached end of life in December 2025.
Pricing
Apache 2.0 has no license fee. The Cassandra, OpenSearch, or ClickHouse cluster behind it is the bill.
Who is Jaeger best for?
Kubernetes teams that prioritize trace-sampling control and can operate a separate storage backend.
4. Grafana Tempo
Grafana Tempo is a high-volume trace store for Grafana teams using object storage instead of an Elasticsearch or Cassandra cluster. It writes Apache Parquet blocks to object storage, and TraceQL reads needed columns. Metrics and logs remain separate data sources in the broader Grafana stack.
Pros
- Object storage tier: The Tempo architecture uses a bucket for trace storage.
- Derived metrics: The metrics generator produces RED metrics and service graphs.
- Multi-protocol ingestion: Tempo ingestion accepts OpenTelemetry, Jaeger, and Zipkin traces.
Cons
- Object-storage requirement: The Parquet architecture requires object storage for columnar Parquet trace blocks.
- Cross-signal discovery: The Grafana data-source model keeps trace discovery within a broader modular stack.
- Separate UI layer: Tempo traces require Grafana as the viewing interface.
Pricing
Apache 2.0 self-hosting is free. Object storage volume and its GET, PUT, and LIST calls are the running cost.
Who is Grafana Tempo best for?
Grafana teams prioritizing object-storage trace retention and accepting a modular discovery workflow.
5. Elastic APM
Elastic APM is an agent-based APM for teams already storing logs in Elasticsearch. The APM Server accepts Elastic language agents, Elastic Distribution of OpenTelemetry (EDOT) agents, and OTLP intake. It then indexes telemetry into Elasticsearch and surfaces traces in Kibana.
Pros
- Shared store: APM telemetry uses the same Elasticsearch cluster and query tools.
- OpenTelemetry intake: OTLP intake arrives beside Elastic’s own agents.
- Language agents: The subscription matrix covers Java, .NET, Go, Ruby, PHP, Python, and Node.js.
Cons
- Index storage: The Elasticsearch data model grows storage with ingest volume.
- Coupled nodes: APM data shares Elasticsearch compute and storage resources.
- Paid analytics: The subscription matrix places anomaly detection and SLO alerting in paid tiers.
Pricing
Elastic’s self-managed Free/Basic tier includes the APM core at no cost. Index growth and additional data nodes remain part of the operating cost.
Who is Elastic APM best for?
Elastic Stack teams prepared to place APM ingest and retention on their existing Elasticsearch capacity model.
6. Zipkin
Zipkin is a tracing system for small teams that prioritize a stable single-process deployment. One process bundles collector, storage, API, and UI on MySQL, Cassandra, Elasticsearch, or OpenSearch, with native B3 propagation. Recent releases focus on dependency and security updates.
Pros
- Mature codebase: Twitter built Zipkin in 2012.
- Instrumentation libraries: The instrumentation libraries cover Java, Python, Go, Ruby, and JavaScript.
- Single process: The server deployment combines collection, storage, API, and UI.
Cons
- Fixed-rate sampling: The server architecture uses a server-side sampling rate.
- No UI authentication: The Zipkin UI has no built-in authentication.
- Partial OTLP: Native ingestion needs the zipkin-otel library.
Pricing
Apache 2.0 has no paid tier. One process and its datastore are the whole bill.
Who is Zipkin best for?
Small teams that prioritize a single-process tracing deployment over native OTLP ingestion and adaptive sampling.
Jaeger vs. Zipkin: Jaeger v2 runs on the OpenTelemetry Collector framework with adaptive and tail-based sampling and ingests OTLP natively; Zipkin runs as one process with a fixed server-side rate and propagates B3.
7. Apache SkyWalking
Apache SkyWalking is an agent-based APM for Java Virtual Machine (JVM) microservice teams that want topology and tracing without changing application code. Agents auto-instrument Java, C#, Node.js, Go, PHP, and Python. The backend accepts OTLP telemetry signals alongside data from SkyWalking agents.
Pros
- Zero-code Java agent: The Java agent instruments without rebuilds or source changes.
- Topology mapping: Horizon renders the service map in a living 3D scene.
- Storage flexibility: Backend configuration supports BanyanDB, Elasticsearch, OpenSearch, MySQL, and PostgreSQL.
Cons
- Agent deployment: The agent architecture places an agent in each instrumented JVM container.
- Version-specific UI: Horizon UI took over in version 11.0.0.
- Manual instrumentation: C++, Rust, and Nginx Lua need manual instrumentation.
Pricing
Apache 2.0 has no license fee. Cost tracks the storage backend and the agent in every container.
Who is Apache SkyWalking best for?
Java microservice teams willing to deploy agents per container to avoid application-source changes.
8. Uptrace
Uptrace is an OpenTelemetry-native backend for small teams wanting traces, metrics, and logs in one place. It stores signals in ClickHouse and pairs PromQL-compatible metric queries with Structured Query Language (SQL)-style trace queries. Self-hosting uses ClickHouse and PostgreSQL, with Docker Compose and Helm deployment paths.
Pros
- Correlated signals: OTLP ingestion places telemetry behind one interface.
- Documented deployment: An official Docker Compose setup is available.
- Defined components: The installation architecture documents ClickHouse and PostgreSQL dependencies.
Cons
- Manual infrastructure: Bare-metal installation requirements include ClickHouse.
- Paid feature tiers: The Uptrace pricing page separates advanced paid features.
- Database dependency: Self-hosting requirements include ClickHouse and PostgreSQL.
Pricing
Self-hosting under the AGPLv3 license is free. Uptrace Cloud lists $0.10 per GB for traces and logs, with volume discounts, while self-hosted teams retain the ClickHouse and PostgreSQL costs.
Who is Uptrace best for?
Small teams prepared to operate ClickHouse and PostgreSQL for a unified OpenTelemetry backend.
9. Pinpoint
Pinpoint is a Java-first APM for teams that need method-level call stacks without modifying code. A -javaagent flag injects bytecode instrumentation at JVM startup. The CallStack, ServerMap, and Inspector views render transactions, topology, and JVM health.
Pros
- Method-level detail: CallStack views reach individual methods.
- Low claimed overhead: The project reports roughly three percent added usage.
- Framework coverage: The Pinpoint repository covers Tomcat, Spring Boot, Kafka, MySQL, and Redis.
Cons
- HBase operations: The Pinpoint architecture requires HBase maintenance, including time to live tuning and major compaction.
- Partial OpenTelemetry: Pinpoint documentation limits OTLP coverage to metrics; tracing uses Pinpoint agents.
- Language skew: The Pinpoint repository routes PHP and Python through a C agent.
Pricing
The software has no license fee under Apache 2.0. The HBase cluster and its tuning are the expense.
Who is Pinpoint best for?
Java teams with HBase experience that prioritize method-level transaction detail over portable trace ingestion.
10. OpenObserve
OpenObserve is a multi-signal observability backend for teams combining logs, metrics, traces, and real user monitoring. It accepts OTLP telemetry and supports single-binary and Helm deployments. Its deployment options include self-hosting and a managed cloud path.
Pros
- Multi-signal scope: The OpenObserve repository covers logs, metrics, traces, and real user monitoring.
- Native OTLP: OTLP ingestion accepts OpenTelemetry data.
- Deployment options: The OpenObserve repository documents self-hosted deployment, including a single binary.
Cons
- Operator-owned backend: Self-hosting the multi-signal backend places storage and upgrades with the operator.
- Capacity coupling: Four signal types share the same OpenObserve deployment and its available capacity.
- OTLP setup required: Collector traffic depends on the documented OTLP ingestion path.
Pricing
The self-hosted software is available under the AGPLv3 license with no license fee. Infrastructure and storage remain the operator’s cost, while managed cloud terms can change.
Who is OpenObserve best for?
Teams seeking one self-hosted backend across four signal types and willing to size shared storage for that scope.
11. Coroot
Coroot is a Kubernetes observability tool for teams seeking coverage without application-level instrumentation. It gathers metrics, traces, logs, and profiling data through OTLP and eBPF. Deployment options include Docker, Helm, an operator, and a DaemonSet.
Pros
- Instrumentation model: The Coroot architecture combines eBPF collection with OTLP logs and traces.
- Kubernetes deployment: Coroot deployment supports a DaemonSet-based agent model.
- Signal coverage: The Coroot overview includes metrics, traces, logs, and profiling.
Cons
- DaemonSet requirement: The Kubernetes deployment depends on DaemonSet support for eBPF collection.
- Fargate limitation: Fargate excludes DaemonSets, so that deployment path is unavailable.
- Kubernetes focus: The Coroot architecture centers its automated coverage on Kubernetes workloads.
Pricing
The self-hosted software uses Apache 2.0 with no license fee. Cluster resources and telemetry storage remain the operator’s cost.
Who is Coroot best for?
Kubernetes teams that can deploy DaemonSets and prioritize eBPF-based coverage over application-agent instrumentation.
How to choose an open source APM tool
Use your existing footprint to narrow the choices before comparing features. Teams already operating Elasticsearch may find Elastic APM operationally familiar; Prometheus users can evaluate the additional components required for Loki and Tempo; JVM teams can compare the agent and storage requirements of SkyWalking and Pinpoint. Then use on-call capacity to settle the rest.
Unified backend vs. modular stack
A unified backend centralizes the query path across logs, metrics, and traces, and your team then owns the database that sits behind it. A modular stack keeps established Prometheus and Grafana components in place and lets each signal scale on its own storage and query components.
Where tracing is the only gap in an otherwise working stack, a trace-only backend fills that gap without disturbing what already works.
Pick the shape that matches how your team already splits ownership of storage and query paths, because changing that shape later is a migration.
OpenTelemetry compatibility
Portability between backends depends on whether the tool accepts OTLP directly. SigNoz, Uptrace, Tempo, Jaeger v2, SkyWalking, and OpenObserve all take logs, metrics, and traces over OTLP. Coroot accepts OTLP logs and traces and pairs them with eBPF-based coverage. Zipkin uses B3 headers that predate W3C TraceContext, and Pinpoint accepts OTLP for metrics only.
Instrumentation on the application side comes from the OpenTelemetry software development kit (SDK) or from zero-code instrumentation using bytecode manipulation or extended Berkeley Packet Filter (eBPF) probes.
Favor an OTLP-native backend if you want the option to switch tools later without reinstrumenting every service.
Storage and operational cost
Storage backend choice decides the self-hosted bill more than the APM tool does. Tempo’s docs use roughly 300 bytes per span as a reference average, so a hypothetical 50,000 spans per second amounts to about 15 MB per second, or roughly 1.3 TB per day before compression. Compression and object-storage tuning move that number substantially.
Jaeger’s ClickHouse backend achieved 8.6x compression on 10 million spans, and a Grafana Tempo discussion reports that S3 cost tuning cut Amazon Simple Storage Service (S3) read costs from $120 to $6 per day.
Model your own span volume against these numbers before shortlisting, because the storage layer, not the tool name, is what shows up on the invoice.
Kubernetes and cloud-native deployment
Most tools in this list ship a supported Kubernetes install path, and the two exceptions constrain your choice. The SigNoz Helm chart, SkyWalking Helm chart, Uptrace Helm charts, and Elastic ECK operator cover Helm and operator deployments, Prometheus installs through kube-prometheus-stack, and Jaeger v2 deploys through the OpenTelemetry Operator.
Zipkin has no official chart, and Fargate does not support DaemonSets, which rules out Coroot’s eBPF agents on that runtime.
If your runtime is Fargate or your team relies on Helm-first workflows, eliminate the tools whose install path does not match before you evaluate anything else.
Enterprise scale considerations
At enterprise scale, retention defaults and storage architecture become the first constraints teams hit. Mimir’s distributed architecture is one option for expanding beyond a single Prometheus server. Cassandra writes extra index records for every service name, operation name, and tag, which is why the Jaeger team recommends OpenSearch for large deployments, and self-hosted observability at this size can require specialist skills.
Retention defaults are short. Self-hosted SigNoz keeps seven days for logs and traces, and Jaeger’s storage example ships with a two-week TTL.
If your compliance or investigation window exceeds those defaults, the tools that scale retention cheaply, such as object-storage backends and columnar stores, move to the top of the shortlist.
Match your open source APM tools to the stack you can run
Self-hosting trades a license fee for recurring engineering work. Inventory the storage, telemetry, and Kubernetes components you already run, verify each shortlisted tool’s OTLP ingestion and deployment path, model retention against span volume, and assign ownership for upgrades, backups, and dashboards. Then run the candidate against production-shaped traffic before committing.
A poor architectural fit costs more than a migration. Storage your team cannot operate turns routine upgrades into reliability work, unclear ownership splits investigations across query languages, and index-first designs pressure teams to shorten retention until the telemetry an on-call engineer needs falls outside the searchable window.
Four situations tend to push teams off a self-hosted stack and onto a managed platform.
- Storage outgrows the managed alternative. Each self-hosted option leaves your team responsible for its trace-storage layer, and that layer is the real comparison against what a managed platform such as Datadog appears to bill per host. Coralogix analyzes data in-stream with Streama© and stores it in your own cloud bucket in open Parquet format, so full telemetry stays available without indexing first.
- Collector tuning becomes a role. Maintenance cost escalates until upgrades slip. Reuse the collectors you already point at SigNoz, Jaeger, Tempo, or Uptrace and switch only the endpoint through OpenTelemetry-native ingestion. Coralogix is 100% OpenTelemetry-native, so services do not need reinstrumentation.
- Root cause is still a human correlating tabs. None of the 11 profiled tools documents Git-correlated root cause analysis. Olly, Coralogix’s autonomous observability agent, cross-references telemetry with Git and surfaced the root cause, blast radius, and line to fix in approximately 4.5 minutes in a demonstrated scenario.
- Auditors want a vendor attestation. Coralogix application performance monitoring runs on a service for which Coralogix lists SOC 2 Type II, PCI DSS, HIPAA, and ISO/IEC 27001 among its compliance programs.
For teams that need managed operations while keeping OpenTelemetry portability, the trade-off worth solving is retaining investigation context on historical telemetry without operating an indexed trace database. Coralogix keeps that context in customer-owned cloud storage so engineers can query historical data without running the database themselves.
Start a free 14-day Coralogix trial with an eight-unit quota and no credit card required, and point your existing OTel collectors at Coralogix to run it alongside your current stack.
Frequently asked questions about open source APM tools
Is SigNoz completely free to use?
The self-hosted Community Edition is free under MIT. SSO/SAML and fine-grained RBAC (Beta) are gated to paid tiers; audit logs and multi-tenancy are listed as coming soon.
Is Elastic APM free and open source?
The self-managed Free/Basic tier includes APM Server, tracing, service maps, and language agents. Binaries remain under the Elastic License.
What free observability tools are available beyond APM?
OpenObserve covers logs, metrics, traces, and real user monitoring (RUM) under AGPLv3. Coroot gathers metrics, traces, logs, and profiling with eBPF.