The short version
- Acceptance rate in the GitHub Copilot admin console is a usage signal, not an ROI signal. Never present it as proof the spend is working.
- Review inactive seats with their owners and check contract terms before treating unused capacity as avoidable cost.
- Measure value downstream against a pre-AI baseline: pull-request throughput, review time, PR size, and change-failure rate, not the keystroke.
- If you cannot state your baseline review time and change-failure rate, that gap is the finding. Instrument first, then decide at renewal.
Start With the Right Question
The question a CFO renews on is not "are engineers using GitHub Copilot?" It is "is this seat spend producing attributable value, measured against how we shipped before?" Those are different questions, and most measurement stops at the first one because it is the easy one.
The same framework applies across coding tools: examine usage, delivery, review, quality, and costs. Available telemetry and billing differ, so adapt the analysis to your plan and work.
One rule holds throughout: anchor every claim to independent research, not to the vendor's marketing. A 10x productivity claim is a slide, not a measurement. One controlled study, an experienced-developer randomized trial from METR in July 2025, found developers were 19% slower with early-2025 AI while believing they were 20% faster. Start from that humility, not from the pitch deck.
The GitHub Copilot Admin Console: Useful, But a Weak Signal
The GitHub Copilot admin console gives you two things quickly: seats assigned versus seats active, and acceptance rate (the share of suggestions developers keep). Pull both. But be clear about what each is worth.
Acceptance rate is a usage-depth signal, not a value signal. A developer can accept a suggestion, then spend twenty minutes verifying and rewriting it. Acceptance counts the keystroke, not the outcome. Treat acceptance rate as evidence that the tool is being touched, never as evidence that it is paying off. Any ROI story built on acceptance rate alone is measuring the wrong end of the pipeline.
- Seats assigned vs active (30-day): the input to your idle-seat math.
- Acceptance rate: usage depth only. Do not present it as ROI.
- Cross-reference with your identity provider (Okta, Entra) for licensed-but-dormant users the console alone may miss.
Review the Cost of Unused Seats
Compare seats you pay for with recent tool activity. Confirm apparently inactive seats with their owners, accounting for leave, seasonal work, and planned demand.
Multiplying inactive seats by unit price shows the cost associated with that capacity, not guaranteed savings. Contract minimums and renewal terms determine whether removing or reallocating seats changes the bill.
The Metrics That Actually Measure Value
Real ROI shows up downstream of the keystroke, in delivery. Measure these against a pre-AI baseline (the same teams, a comparable window before rollout), pulled from GitHub Insights or your engineering-intelligence tool:
- Pull-request throughput: merged PRs per developer per week, now versus your pre-GitHub Copilot baseline. Treat the count as one delivery signal, checking task size, quality, and changes in how work is split.
- Pull-request size trend: rising PR size means larger, harder-to-review batches. Growing size is a warning, not a win.
- Review and merge time: check whether more generated code is increasing review effort or waiting time. Faros AI, a measurement vendor, reported review time rising 91% in one 10,000-developer study. Treat this as vendor telemetry from a particular sample, not a forecast for your team.
- Change-failure and revert rate: the quality check. DORA's 2025 research linked AI adoption to higher throughput while identifying continuing delivery-quality challenges. Speed that ships more incidents is not ROI.
- Spend versus value by team: include licenses, usage charges, training, review, and maintenance. Explain how you value delivery or quality improvements and where attribution remains uncertain. PR counts and quality deltas are inputs to that estimate, not its monetary numerator.
The Perception Gap Is Part of the Finding
Ask your team how much faster they feel with GitHub Copilot and how much of that speedup survives review and rework. Compare those answers with measured outcomes; they may agree, diverge, or remain uncertain.
This is not a knock on your engineers. The METR trial found developers felt +20% while measuring -19%, a roughly 40-point swing between perception and reality. Perceived productivity is an unreliable instrument because the verification and integration cost of AI-generated code is invisible in the moment and expensive in aggregate. When you present measured deltas next to felt ones, the contrast is disarming and it reframes the whole conversation away from vibes and toward numbers.
If You Cannot State a Baseline, That Is the Finding
Many teams run this exercise and discover they cannot answer basic questions from memory: What was our review time before GitHub Copilot? Our change-failure rate? Our PR throughput per developer? If those numbers do not exist, no honest ROI verdict is possible yet, for any tool.
That is not a failed measurement. It is the most important result you will get. The recommendation writes itself: instrument the handful of metrics above before scaling spend further, then re-measure against a real baseline in a quarter. A team that cannot state its baseline is not in a position to renew or cut with confidence, and naming that gap plainly is more valuable than any manufactured number.
Consider right-sizing, keeping or expanding, changing the workflow, or cutting. Different teams may need different actions; flat throughput and rising review time warrant investigation before choosing a cause. McKinsey found workflow redesign associated with self-reported EBIT impact, which does not establish causation.
Sources
Original research and publisher reports, with study limits and commercial sources identified.
- METR: experienced developers and early-2025 AI tools
Randomized trial; the 19% slowdown and perceived 20% speedup refer to this sample and period.
- METR: February 2026 productivity update
Follow-up with results consistent with speedups, but selection and timing biases prevent a reliable current estimate.
- DORA: 2025 State of AI-assisted Software Development
Google-led observational research; throughput and delivery-quality findings are associations.
- Faros AI: AI productivity research in 2025
Measurement-vendor telemetry from 10,000 developers; reports 91% longer review time.
- McKinsey: how organizations are rewiring to capture value
Survey associations with self-reported EBIT impact; does not establish that workflow redesign causes gains.
See where you stand.
Use the free self-check to identify which adoption, delivery, and review metrics you can already measure, and which ones need a baseline.
Questions leaders ask
Is GitHub Copilot's acceptance rate a good measure of ROI?
No. Acceptance rate tells you a suggestion was kept, not that it saved net time or shipped safely. A developer can accept a suggestion and then spend significant time verifying and reworking it. Use acceptance rate from the GitHub Copilot admin console as a usage-depth signal only, and measure ROI on downstream delivery metrics like pull-request throughput, review time, and change-failure rate against a pre-AI baseline.
What metrics actually measure GitHub Copilot ROI?
Idle seats (licensed vs active), pull-request throughput versus your pre-GitHub Copilot baseline, pull-request size trend, review and merge time, change-failure or revert rate, and spend-vs-value by team. The per-team cross-tab of seat cost against measured delivery gains is the one finance decisions are made on. Acceptance rate and felt speedup are context, not the answer.
Where does the GitHub Copilot data live?
Seats assigned versus active and acceptance rate come from the GitHub Copilot admin console, cross-referenced with your identity provider for dormant licenses. Throughput, PR size, review time, and revert rate come from GitHub Insights or an engineering-intelligence tool. Seat count and unit price come from your finance or procurement contact. No source-code access is required to answer the ROI question.
Why do developers feel faster with GitHub Copilot than the data shows?
Because perceived productivity is an unreliable instrument. The visible win (code appears instantly) is obvious, while the hidden cost (verifying, integrating, and reworking AI-generated code, plus slower downstream review) is not felt in the moment. METR's 2025 randomized trial found experienced developers felt about 20% faster while measuring roughly 19% slower, a gap worth checking for rather than assuming applies to your team.
What if we do not have a baseline to compare against?
State what is missing and establish a comparable baseline for delivery, review, quality, and costs. Existing records may support a partial retrospective comparison. If renewal comes first, make the uncertainty explicit and weigh a limited commitment or pilot against contract deadlines and switching costs.