Use cases
APM v2 answers a different question depending on what you are doing and who you are. Find your task in Common scenarios, follow a walkthrough, or jump to your role below. Each entry links to the feature behind it.
Common scenarios
- Trace a request end to end: follow a request across services in the Traces tab, see the topology in the Service Map, and inspect upstream and downstream calls in Dependencies.
- Diagnose a latency spike: read the RED charts on the entity Overview, find the slow route in Transactions, and rule out the runtime or host in the Runtime and Infrastructure tabs.
- Investigate an error surge: group failures by status code in the Errors tab, then pivot to the Traces and Logs behind them.
- Monitor service health and SLOs: track health across the fleet in the catalog, and manage SLOs, alerts, and cases in the Monitoring tab. See service health and Service SLOs.
- Correlate a change with a regression: overlay deployments and configuration changes with annotations, and compare current performance against an earlier period.
- Monitor database performance: switch the catalog to Databases to compare query volume, latency, and failures, then use a service's Dependencies to see which services drive the load.
- Control span-metric cost: keep cardinality in check with cardinality limiting and sampling.
Walkthroughs
These worked examples take a task from symptom to cause, naming the exact controls to use at each step.
Diagnose a latency spike
You'll need: a service reporting to APM v2, and a time range covering the spike. Time: about 5 minutes. You'll end with: the transaction, dependency, or host responsible, and the traces that prove it.
- In the Coralogix toolbar, select APM, then open Services.
- Sort by P95 latency and select the service that spiked. Its Overview tab opens on the RED cards.
- On the Latency card, select P95 and P99 together. If P99 moved but P95 did not, this is a tail problem - a small share of slow requests - not a service-wide regression.
- Set Compare to → previous consecutive period. Each card now shows a delta against the equivalent window, so you can tell a spike from normal daily shape.
- Read the Avg latency by dependency card. If one backend dominates, skip to step 7.
- Open Transactions and sort by Time consuming, not by latency - the transaction consuming the most total time is the one worth fixing. Open it to see its segments and spans.
- Rule out the layers beneath the code: Runtime for GC pressure and thread contention on JVM services, Infrastructure for CPU, memory, and network pressure on its hosts.
- Confirm it: open Traces, filter to the slow transaction, and open a trace from inside the spike. The longest span is your answer.
- Find out what changed: turn on Annotations and look for a deployment or configuration change at the start of the spike.
If latency is high but no transaction stands out, the load is spread - check Infrastructure for a shared-resource limit before looking further into the code.
Next: set a latency threshold that catches this earlier - see Service health.
Investigate an error surge
You'll need: a service with a rising error rate, and a time range covering the surge. Time: about 5 minutes. You'll end with: the failing operation, the status code behind it, and the traces and logs that explain it.
- In the Coralogix toolbar, select APM, then open Services.
- Sort by Error rate and select the service that surged. Its Overview tab opens on the RED cards.
- Set Compare to → previous consecutive period to separate a real surge from normal error noise.
- Open the Errors tab. Failures group by status code for services (or by operation for databases) - open the largest group.
- Read Failing transactions to see which requests carry the errors, and use the error-class filter (Server errors / Client errors) to tell a broken dependency from bad input.
- Open an error occurrence to jump to its span, then its full trace - the failing span shows the exception and where it was thrown.
- Pivot to the Logs tab for the log lines around the failure.
- Find out what changed: turn on Annotations and look for a deployment at the start of the surge.
Next: turn the recurring failure into an alert from the Monitoring tab.
Find the slow database behind a slow service
You'll need: a service whose latency is driven by a backend, and a time range covering the slowdown. Time: about 5 minutes. You'll end with: the database operation responsible and the query spans that prove it.
- From the slow service's Overview tab, read the Avg latency by dependency card - if a database dominates the time per request, that's your lead.
- Open the service's Dependencies tab and switch to the Databases view. Rank by Total time spent to find the database the service waits on most.
- Open that database's drilldown drawer and read its Root transactions breakdown to see which of the service's transactions drive the load.
- Switch the catalog to Databases and open the database as an entity. Its Operations tab ranks operations by time consumed.
- Open the slowest operation to inspect its query spans and the traces they belong to - the query text and duration are on the span.
Next: compare the database against an earlier period to confirm the regression, from the Databases catalog.
For on-call responders
You arrive mid-incident, usually from a page rather than the catalog, and need to get from alert to cause fast.
- Start from the alert or case: an entity's Monitoring tab shows the alerts and cases attached to it, and a firing alert has already opened a case.
- Confirm the symptom: the entity Overview RED cards, with Compare to on, show whether latency or errors moved and by how much.
- Find the cause: follow Diagnose a latency spike or Investigate an error surge above.
- Read the blast radius: the Service Map and Dependencies show everything upstream and downstream of the failing entity.
- Line up the trigger: annotations overlay the deployment or change that started it.
For application engineers
You own a service and need to know why it is slow or failing, and where in the code to look.
- Start at your service: open it from the catalog to see its RED metrics and health at a glance.
- Find the failing requests: in the Errors tab, group errors by status code, open the top group, and jump to the errored span.
- Isolate the slow path: in the Transactions tab, find the slowest route or operation and drill into its segments and spans.
- See the request end to end: pivot to the Traces tab for the full trace, and the Logs tab for the log lines around it.
- Rule out the runtime: for JVM services, the Runtime tab surfaces memory leaks, GC pressure, and thread contention.
- Check downstream: the Dependencies tab shows whether a database or external call is the real cause.
For DevOps and SRE
You keep a fleet of services healthy, respond to incidents, and stop regressions before they reach users.
- Scan fleet health: the catalog Health view and service health map show which services are critical, warning, healthy, or unmonitored. Filter by environment to focus on production.
- Define reliability targets: track SLOs and error budgets in the Monitoring tab. See Service SLOs.
- Respond to issues: alerts and cases attached to an entity live in the same Monitoring tab; create alerts straight from a widget. See APM monitoring with alerts.
- Correlate with deployments: annotations overlay deployments and configuration changes on the charts so you can line up a spike with a release.
- Separate infra from application problems: the Infrastructure tab shows CPU, memory, and network pressure on the hosts an entity runs on.
- Understand blast radius: the Service Map and Dependencies show what a failing service affects.
- Catch regressions: compare current performance against a previous period from the catalog.
For engineering managers
You track the reliability posture of your team's services and decide where to invest.
- See ownership at a glance: the catalog Ownership view adds owner, alerts, cases, and SLO status columns. Filter by team to see only your services.
- Watch SLO compliance: Service SLOs and their error budgets show which services are meeting their targets and which are burning budget.
- Spot chronic problems: service health trends surface the services that sit in warning or critical over time.
- Measure user satisfaction: the Apdex score turns latency into a single satisfaction signal per service.
For database engineers
You monitor database performance and the services that put load on your databases.
- Monitor databases as entities: select Databases in the catalog to compare query volume, latency, and failures per database.
- Break performance down by operation: a database's Operations tab in the drilldown shows query throughput, latency, and errors per operation and table, plus how many services call each one.
- Investigate database errors: a database's drilldown surfaces its error groups and recent traces.
- Trace load back to its source: a service's Dependencies tab shows which services query a database and how those queries perform.
- See the topology: the Service Map shows how your services connect to their databases.
Next steps
Instrument your first service in the APM onboarding tutorial, then explore it in the catalog.