Skip to main content

Use cases

APM v2 answers a different question depending on what you are doing and who you are. Find your task in Common scenarios, follow a walkthrough, or jump to your role below. Each entry links to the feature behind it.

Common scenarios​

  • Trace a request end to end: follow a request across services in the Traces tab, see the topology in the Service Map, and inspect upstream and downstream calls in Dependencies.
  • Diagnose a latency spike: read the RED charts on the entity Overview, find the slow route in Transactions, and rule out the runtime or host in the Runtime and Infrastructure tabs.
  • Investigate an error surge: group failures by status code in the Errors tab, then pivot to the Traces and Logs behind them.
  • Monitor service health and SLOs: track health across the fleet in the catalog, and manage SLOs, alerts, and cases in the Monitoring tab. See service health and Service SLOs.
  • Correlate a change with a regression: overlay deployments and configuration changes with annotations, and compare current performance against an earlier period.
  • Monitor database performance: switch the catalog to Databases to compare query volume, latency, and failures, then use a service's Dependencies to see which services drive the load.
  • Control span-metric cost: keep cardinality in check with cardinality limiting and sampling.

Walkthroughs​

These worked examples take a task from symptom to cause, naming the exact controls to use at each step.

Diagnose a latency spike​

You'll need: a service reporting to APM v2, and a time range covering the spike. Time: about 5 minutes. You'll end with: the transaction, dependency, or host responsible, and the traces that prove it.

  1. In the Coralogix toolbar, select APM, then open Services.
  2. Sort by P95 latency and select the service that spiked. Its Overview tab opens on the RED cards.
  3. On the Latency card, select P95 and P99 together. If P99 moved but P95 did not, this is a tail problem - a small share of slow requests - not a service-wide regression.
  4. Set Compare to → previous consecutive period. Each card now shows a delta against the equivalent window, so you can tell a spike from normal daily shape.
  5. Read the Avg latency by dependency card. If one backend dominates, skip to step 7.
  6. Open Transactions and sort by Time consuming, not by latency - the transaction consuming the most total time is the one worth fixing. Open it to see its segments and spans.
  7. Rule out the layers beneath the code: Runtime for GC pressure and thread contention on JVM services, Infrastructure for CPU, memory, and network pressure on its hosts.
  8. Confirm it: open Traces, filter to the slow transaction, and open a trace from inside the spike. The longest span is your answer.
  9. Find out what changed: turn on Annotations and look for a deployment or configuration change at the start of the spike.

If latency is high but no transaction stands out, the load is spread - check Infrastructure for a shared-resource limit before looking further into the code.

Next: set a latency threshold that catches this earlier - see Service health.

Investigate an error surge​

You'll need: a service with a rising error rate, and a time range covering the surge. Time: about 5 minutes. You'll end with: the failing operation, the status code behind it, and the traces and logs that explain it.

  1. In the Coralogix toolbar, select APM, then open Services.
  2. Sort by Error rate and select the service that surged. Its Overview tab opens on the RED cards.
  3. Set Compare to → previous consecutive period to separate a real surge from normal error noise.
  4. Open the Errors tab. Failures group by status code for services (or by operation for databases) - open the largest group.
  5. Read Failing transactions to see which requests carry the errors, and use the error-class filter (Server errors / Client errors) to tell a broken dependency from bad input.
  6. Open an error occurrence to jump to its span, then its full trace - the failing span shows the exception and where it was thrown.
  7. Pivot to the Logs tab for the log lines around the failure.
  8. Find out what changed: turn on Annotations and look for a deployment at the start of the surge.

Next: turn the recurring failure into an alert from the Monitoring tab.

Find the slow database behind a slow service​

You'll need: a service whose latency is driven by a backend, and a time range covering the slowdown. Time: about 5 minutes. You'll end with: the database operation responsible and the query spans that prove it.

  1. From the slow service's Overview tab, read the Avg latency by dependency card - if a database dominates the time per request, that's your lead.
  2. Open the service's Dependencies tab and switch to the Databases view. Rank by Total time spent to find the database the service waits on most.
  3. Open that database's drilldown drawer and read its Root transactions breakdown to see which of the service's transactions drive the load.
  4. Switch the catalog to Databases and open the database as an entity. Its Operations tab ranks operations by time consumed.
  5. Open the slowest operation to inspect its query spans and the traces they belong to - the query text and duration are on the span.

Next: compare the database against an earlier period to confirm the regression, from the Databases catalog.

For on-call responders​

You arrive mid-incident, usually from a page rather than the catalog, and need to get from alert to cause fast.

  • Start from the alert or case: an entity's Monitoring tab shows the alerts and cases attached to it, and a firing alert has already opened a case.
  • Confirm the symptom: the entity Overview RED cards, with Compare to on, show whether latency or errors moved and by how much.
  • Find the cause: follow Diagnose a latency spike or Investigate an error surge above.
  • Read the blast radius: the Service Map and Dependencies show everything upstream and downstream of the failing entity.
  • Line up the trigger: annotations overlay the deployment or change that started it.

For application engineers​

You own a service and need to know why it is slow or failing, and where in the code to look.

  • Start at your service: open it from the catalog to see its RED metrics and health at a glance.
  • Find the failing requests: in the Errors tab, group errors by status code, open the top group, and jump to the errored span.
  • Isolate the slow path: in the Transactions tab, find the slowest route or operation and drill into its segments and spans.
  • See the request end to end: pivot to the Traces tab for the full trace, and the Logs tab for the log lines around it.
  • Rule out the runtime: for JVM services, the Runtime tab surfaces memory leaks, GC pressure, and thread contention.
  • Check downstream: the Dependencies tab shows whether a database or external call is the real cause.

For DevOps and SRE​

You keep a fleet of services healthy, respond to incidents, and stop regressions before they reach users.

  • Scan fleet health: the catalog Health view and service health map show which services are critical, warning, healthy, or unmonitored. Filter by environment to focus on production.
  • Define reliability targets: track SLOs and error budgets in the Monitoring tab. See Service SLOs.
  • Respond to issues: alerts and cases attached to an entity live in the same Monitoring tab; create alerts straight from a widget. See APM monitoring with alerts.
  • Correlate with deployments: annotations overlay deployments and configuration changes on the charts so you can line up a spike with a release.
  • Separate infra from application problems: the Infrastructure tab shows CPU, memory, and network pressure on the hosts an entity runs on.
  • Understand blast radius: the Service Map and Dependencies show what a failing service affects.
  • Catch regressions: compare current performance against a previous period from the catalog.

For engineering managers​

You track the reliability posture of your team's services and decide where to invest.

  • See ownership at a glance: the catalog Ownership view adds owner, alerts, cases, and SLO status columns. Filter by team to see only your services.
  • Watch SLO compliance: Service SLOs and their error budgets show which services are meeting their targets and which are burning budget.
  • Spot chronic problems: service health trends surface the services that sit in warning or critical over time.
  • Measure user satisfaction: the Apdex score turns latency into a single satisfaction signal per service.

For database engineers​

You monitor database performance and the services that put load on your databases.

  • Monitor databases as entities: select Databases in the catalog to compare query volume, latency, and failures per database.
  • Break performance down by operation: a database's Operations tab in the drilldown shows query throughput, latency, and errors per operation and table, plus how many services call each one.
  • Investigate database errors: a database's drilldown surfaces its error groups and recent traces.
  • Trace load back to its source: a service's Dependencies tab shows which services query a database and how those queries perform.
  • See the topology: the Service Map shows how your services connect to their databases.

Next steps​

Instrument your first service in the APM onboarding tutorial, then explore it in the catalog.

Last updated on