All guides

How to Measure Claude Code ROI (When You Can't Tell If the Spend Is Working)

You bought Claude Code seats, the team says it feels faster, and now finance wants a number before renewal. Here is how to measure whether the spend is actually working, using signals you can trust instead of a vendor's acceptance rate.

In this guide

The short version

  • Acceptance rate is a usage signal, not proof of return. METR's early-2025 sample illustrates why perceived speed should be checked against measured outcomes.
  • Six signals to examine: idle seats, PR throughput versus baseline, review-time trend, change-failure/revert rate, PR size, and spend versus value by team.
  • Combine usage information with delivery, quality, and cost records to assess value. Adapt metric definitions and costs to your tool and plan.
  • If your team cannot state its pre-AI baseline for review time and change-failure rate, that gap is the finding: instrument first, then scale, then renew on evidence.

Start with the question renewal is really asking

The question is not "do engineers like Claude Code?" The question your CFO is asking is narrower and harder: for the dollars you are spending on seats, what attributable delivery gain came back, and by which teams?

That is a measurement problem, not a tool problem. And it has a specific starting point: your pre-AI baseline. If you cannot state, from memory or from a dashboard, your review time and change-failure rate from before the rollout, you cannot claim a delta. That gap is not a failure of the engagement. It is the finding.

The same questions apply across coding tools: who uses it, what changes in delivery and quality, and what does it cost? Adapt the collection to your tool, plan, and kind of work; available telemetry and billing differ.

Why acceptance rate is the trap, not the metric

The number every tool shows you first is a usage or acceptance signal: how often suggestions get taken, how active a seat is. It is comforting and it is nearly worthless as an ROI proxy. High acceptance tells you people are typing tab, not that your organization ships more or ships safer.

The reason is the productivity paradox, and it is measured, not anecdotal. In a 2025 randomized controlled trial, METR found experienced open-source developers were about 19% slower with early-2025 AI tools, while those same developers believed they were about 20% faster, even after the fact. Perceived speed and measured speed pointed in opposite directions.

So the first discipline is to distrust the feeling and the acceptance chart alike, and to instrument the outcomes those charts are supposed to predict.

The metrics that actually decide ROI

Anchor on outcomes, then layer in AI-specific signals. These six measures are starting points; a monetary ROI estimate also needs a defensible way to value benefits and account for total costs:

  • Idle seats: licensed seats versus 30-day active seats. This is a concrete place to look for avoidable license costs. A dormant seat is a candidate for review: check leave, seasonal work, planned demand, and contract terms before removing it.
  • Pull-request throughput versus a pre-AI baseline: merged PRs per developer per week, now compared to before. This is one delivery signal; changes in task size, staffing, and quality can also change the count.
  • Review-time trend: distinguish waiting time from active review effort where possible. Faros AI, a measurement vendor, reported 91% longer reviews in one study and roughly five times the median in a later sample. These observations warrant investigation; they do not establish the effect in your team.
  • Change-failure and revert rate: does more code mean more rollbacks? Google-led DORA research links AI adoption to higher throughput while identifying continuing delivery-quality challenges, so watch this one closely.
  • Spend versus attributable value, per team: include licenses, usage charges, training, review, and maintenance. Explain how you value benefits and what remains uncertain. A throughput delta alone is not a dollar return.
  • PR size trend: larger changes alongside longer reviews are a reason to inspect batch size and review capacity. Check task mix and generated files before attributing the pattern to AI.

Where Claude Code's numbers live (and why the tool is the easy part)

For adoption, review the admin and usage data available under your Claude Code plan. Reconcile user activity with provisioned access and the invoice, including usage charges and renewal terms. Identity-provider sign-ins can help identify accounts to review, but do not establish activity inside the tool.

For delivery, use version-control and CI/CD data where available: completed work, pull-request size, review time, deployment frequency, and reverts. Join these with incident records and costs. Check how each system defines its metrics and what it does not capture.

The framework can carry across tools, but definitions and costs still need checking. Compare similar work over comparable periods, record other changes, and explain what the evidence can and cannot attribute to the tool.

Measure the perception gap on purpose

The most persuasive slide in any ROI readout is not a throughput chart. It is the gap between what the team believes and what the data shows.

Run a two-minute pulse across the team: "How much faster do you feel with AI, as a percentage?" and "How much of that speedup survives review and rework?" Then put the answers next to your measured throughput and review-time trends.

When the felt number is +30% and the measured number is flat or negative, the gap tells you to investigate. Drafting may be faster while review, verification, and integration take longer; the perceived speedup may also overstate the gain. Measure each stage before deciding what to change. One independent academic study offers a possible explanation: Carnegie Mellon University, across 807 repositories, found Cursor adoption was associated with roughly 41% higher code complexity. Separately, CodeRabbit, an AI code-review vendor, reviewed 470 open-source pull requests and found the AI-coauthored ones carried about 1.7x more issues. These studies support checking review load and code quality; they do not establish what happened on your team.

If you cannot state your baseline, that is the answer

A useful readout distinguishes measured changes, estimated benefits and costs, and unresolved questions. You may have enough evidence to keep, right-size, or expand, or need a limited pilot or better instrumentation before a larger commitment.

Missing data is a useful finding. Start recording delivery, review, quality, and seat activity, and state which parts of the decision remain uncertain. Contract timing and switching costs still matter while you gather evidence.

Keep the source and its limits beside each benchmark: METR's controlled study, the academic Cursor preprint, Google-led DORA research, and labeled vendor telemetry answer different questions. None substitutes for evidence about your own team.

Sources

Original research and publisher reports, with study limits and commercial sources identified.

  1. METR: experienced developers and early-2025 AI tools

    Randomized trial; the 19% slowdown and perceived 20% speedup refer to this sample and period.

  2. METR: February 2026 productivity update

    Follow-up with results consistent with speedups, but selection and timing biases prevent a reliable current estimate.

  3. DORA: 2025 State of AI-assisted Software Development

    Google-led observational research; throughput and delivery-quality findings are associations.

  4. Cursor adoption across 807 open-source repositories

    Academic preprint; the roughly 41% complexity increase is an observed association.

  5. Faros AI: AI productivity research in 2025

    Measurement-vendor telemetry from 10,000 developers; reports 91% longer review time.

  6. Faros AI: AI Acceleration Whiplash

    Measurement-vendor telemetry from 22,000 developers; reports roughly five times the median review time.

  7. CodeRabbit: AI vs. human code-generation report

    Code-review vendor study of 470 pull requests; reports about 1.7 times as many issues in AI-coauthored code.

  8. McKinsey: how organizations are rewiring to capture value

    Survey associations with self-reported EBIT impact; does not establish that workflow redesign causes gains.

See where you stand.

Use the free self-check to identify which adoption, delivery, and review metrics you can already measure, and which ones need a baseline.

Questions leaders ask

Isn't Claude Code's acceptance rate a good enough ROI signal?

No. Acceptance and usage rates tell you people are using the tool, not that your organization ships more or ships safer. METR's 2025 randomized trial found developers who felt about 20% faster were measured about 19% slower, so a usage chart can move in the opposite direction from real delivery. Measure outcomes (throughput versus baseline, review time, change-failure rate), and use the admin telemetry only for the adoption and idle-seat layer.

We never recorded a baseline before rolling out Claude Code. Can we still measure ROI?

Partially, and the gap itself is the most valuable finding. Without a pre-AI baseline you cannot claim a clean throughput delta, but you can still surface idle-seat waste, current review-time and change-failure trends, and the perception gap. The honest recommendation becomes: instrument these four signals now, before scaling further, so the next renewal decision runs on evidence instead of feel.

Where should we check for avoidable costs first?

Start by comparing licensed seats with 30-day activity and checking the exceptions with team leads. Leave, seasonal work, and upcoming projects can explain inactivity. Contract terms determine whether removing or reallocating a seat saves money; an audit does not guarantee a saving.

Is this method specific to Claude Code?

The framework applies across tools: examine adoption, delivery, review, quality, and total costs. Available telemetry, metric definitions, usage charges, and the work itself differ, so adapt the analysis.

Where do the ROI benchmark numbers come from, and can I trust them?

Anchor to independent research: METR's randomized trial on perceived versus measured speed, Google-led DORA research on throughput and delivery-quality challenges, Carnegie Mellon University on code complexity, and McKinsey on workflow redesign as the top correlate of impact. Measurement-vendor figures like those from Faros or DX are useful but self-interested, so label them as vendor telemetry and never present them as independent. Treat all of it as directional, since the evidence is young and moves fast.