Back

AI Impact Measurement: The Three Levels Explained

AI Impact Measurement: The Three Levels Explained

AI impact measurement tells you what your AI spend is buying. Most engineering teams can report how many tokens they burned last month, but very few can say how many of those tokens became merged code, or what that code did once customers used it.

This guide splits AI impact measurement into three levels: tokenomics, code impact and production impact. Each level answers a harder question and needs the data from the level beneath it.

What Is AI Impact Measurement?

AI impact measurement is the practice of tracing AI activity, from coding agent sessions to production AI workloads, through to its effect on cost, engineering output and the experience of real users. It joins usage data with delivery data and production telemetry, so you can follow a single model call to the pull request it produced and the error rate after that code shipped.

It covers two kinds of AI. Coding agents such as Claude Code, Cursor, Codex and GitHub Copilot write your software, while production AI workloads like chat assistants and agentic workflows run inside it. The same three levels apply to both.

The Three Levels of AI Impact Measurement

The three levels of AI impact measurement: tokenomics asks what you spent, code impact asks what it produced, production impact asks what it changed

Most organizations stop at tokenomics, since every provider’s billing console reports token counts for free. Levels two and three need data from outside the provider: your repositories, your CI pipeline and your production telemetry.

Level 1: Tokenomics

Tokenomics counts token burn and turns it into spend, so you can see where the money goes.

Every model call reports input, output and cached tokens, and every provider publishes a price per million. Multiply them out, then group the total by the dimensions that explain it:

  • Model, because a frontier model can cost many times more per token than a smaller one
  • Provider, since Anthropic, OpenAI and Google each bill and discount caching differently
  • Effort or reasoning level, where higher settings burn more output tokens on the same task
  • Tool, as Claude Code, Cursor and Copilot sessions each have their own usage profile
  • Team, repository and developer, for the moment finance asks who spent what

Tokenomics gives you a budget line, spike alerts and a first read on adoption. It can’t tell you whether the spend was worth it, because a developer who burns $400 in a week might be shipping half your roadmap. Or looping on one failing test.

Level 2: Code Impact

Code impact measures the work itself: what each agent session set out to do, and what came out of it.

Start with sessions. Each one has an intent, like fixing a bug, adding an endpoint or refactoring a module, and it either lands as a well-defined pull request or it doesn’t. The core code impact metrics follow that session through review:

  • Session to PR rate, the share of sessions that produced a reviewable pull request
  • Merge rate, the share of those pull requests that merged
  • Time to PR, from the first prompt to an open pull request
  • Rework, meaning follow-up commits, reverts and extra review rounds on AI-authored changes
  • Cost per merged PR, spend divided by merged output, and the first figure that ties cost to results

Two teams with identical token bills can sit at $3 and $30 per merged PR.

Code impact data comes from your repositories, so every session has to be attributed to the repository it touched. Attribution also separates company code from personal or unmanaged projects still billed to a company account.

Level 3: Production Impact

Production impact connects AI activity to what happened in your production systems and to your users.

For coding agents, follow each merged pull request through deployment and compare error rates and endpoint latency on the services it changed, before and after the release. A pull request that merged in 20 minutes and caused a rollback two hours later looks like a success at level two. Only level three shows the rollback.

For production AI workloads, apply the same idea to every request. Trace a chat turn or an agent run through its model calls, tool calls and guardrail decisions. Then check what the user did next. An answer that ends the conversation is a different outcome from one the user immediately rephrases, even when both cost the same.

An AI change measured from agent session to pull request, release and production, with the level that measures each stage

Session data, repository events and production telemetry usually sit in three tools with no shared identifier. Production impact measurement needs them matched, so any trace can be tied to the commit that shipped it and the session that wrote it. OpenTelemetry’s Gen AI semantic conventions give model calls standard attributes to match on.

What Good AI Impact Measurement Looks Like

A mature program measures all three levels continuously and can move between them in one investigation:

  1. Every agent and workload reports telemetry. Coding agents and production AI both export usage over OpenTelemetry, with no gaps by team or tool.
  2. Spend is attributed. Every dollar maps to a model, a repository and a person.
  3. Output sits next to spend. Merge rate, time to PR and rework share a dashboard with cost.
  4. Production is in the loop. You can trace a spike in cost or errors back to the session behind it.
  5. Delivery stays healthy. Standard DORA metrics such as change failure rate should hold steady or improve as AI adoption grows.

If you can answer “what did this model cost, what did it ship, and what did that change do in production” for any week you choose, the program is working.

Frequently Asked Questions

What should I measure first?

Start with tokenomics, grouped by model, tool and repository. It needs only provider usage data, and repository attribution sets you up for code impact later.

How is AI impact measurement different from AI observability?

AI observability watches AI systems in production: latency, errors, cost and output quality per request. AI impact measurement uses that data, along with coding agent and repository data, to judge whether the AI investment as a whole is paying off.

Can AI impact measurement prove ROI?

It gets you the inputs. Cost per merged PR, time to PR and production error rates on AI-authored changes give finance and engineering leaders figures they can track quarter over quarter. To see how Coralogix connects all three levels in one place, visit AI Impact Measurement.

On this page