The short version
- Review inactive seats with their owners and check contract terms before removing or reallocating them.
- Judge value on measured pull-request throughput, review time, and change-failure rate versus your pre-AI baseline, not on how fast the team feels.
- Check perception against outcomes. METR found a gap in its early-2025 sample; do not assume that result applies to current tools or your team.
- Consider keeping, renegotiating, improving the workflow, or cutting. The right action can differ by team; investigate causes before choosing.
- If you cannot state your baseline, that gap is the finding. Instrument first, then decide.
The renewal question to answer
Your AI coding seats are up for renewal. Finance wants to know if the spend is producing a return. Your engineers swear the tools make them faster. And when you try to put a number behind either claim, you find you do not actually have one.
That gap is not a failure. It is the single most important finding of this whole exercise. If your team cannot state its own baseline for pull-request throughput, review time, and change-failure rate, then you are not renewing on evidence. You are renewing on vibes. The good news: the decision is tractable, the data mostly already exists, and the method is the same no matter which tool you bought, whether that is GitHub Copilot, Cursor, or Claude Code.
This guide considers keeping, renegotiating, improving the surrounding workflow, or cutting. The choice may differ by team, and incomplete evidence can justify a limited pilot or temporary renewal.
Start with the cheapest win: idle seats
Before you measure anything sophisticated, count seats. Not seats licensed. Seats actually active in the last 30 days.
Compare the activity data available in your plan with your invoice and identity provider. Ask team leads about leave, seasonal usage, onboarding, and upcoming projects before treating a seat as unnecessary. Review contract minimums and renewal terms: reduced seat counts do not always produce immediate savings.
- Pull licensed seat count and unit price from your actual invoice or contract, broken out by team.
- Pull 30-day active users from the tool's admin console.
- Review each inactive seat with its owner, then decide whether to keep, reallocate, or remove it under your contract.
The metrics that actually decide it
Idle seats tell you about waste. To judge value, you need to compare delivery today against your pre-AI baseline. Adoption depth signals like acceptance rate are weak and easy to over-trust, so treat them as context, not proof. The signals that matter live in your version-control analytics (GitHub or GitLab Insights) and your incident tooling.
- Pull-request throughput (merged PRs per developer per week) versus the pre-AI baseline. Check task size and quality alongside the count; more PRs alone do not establish return.
- Pull-request size trend. Larger changes can make review harder; inspect task mix, generated files, and review practices before drawing a conclusion.
- Review and merge time trend. Check for downstream waiting and additional review effort. Faros AI (a measurement vendor, citing its own telemetry) reported review time rising 91% in one 10,000-developer study and roughly 5x in a later 22,000-developer study as adoption deepened.
- Change-failure and revert rate. DORA's 2025 research linked AI adoption to higher throughput while identifying continuing delivery-quality challenges.
- Spend versus attributable value, per team. Include relevant costs, explain how benefits are valued, and make uncertainty and other changes in the period visible.
The perception gap you have to confront
Your engineers' sense of how much faster they are deserves to be measured alongside actual outcomes. A self-reported speedup alone cannot establish the return.
In a randomized controlled trial published in July 2025, METR found experienced open-source developers were 19% slower with early-2025 AI tooling, while those same developers believed they were about 20% faster, even after the fact. Felt plus twenty, measured minus nineteen. The overhead hides in verification and integration: reading, testing, and repairing code you did not write yourself.
The study does not show that current tools generally slow developers down. Team feedback is useful evidence about the experience of using a tool, but it needs to be checked against delivery, quality, and cost data.
The four-way decision
Use adoption, delivery, review effort, quality, and costs to consider four possible actions. Different teams can need different choices, and a short pilot or temporary renewal can be appropriate when evidence is incomplete.
- Active seats well below licensed: review the exceptions, then negotiate, reallocate, or right-size where contract terms allow.
- Adoption healthy, worthwhile benefits, and acceptable quality and costs: consider keeping or expanding the tool. State the evidence and uncertainty behind the decision.
- Adoption is healthy but delivery is flat while review time or defects rise: investigate review capacity, task mix, integration, and tool fit. Test the suspected bottleneck before deciding whether to change the process, the tool, or both.
- No adoption, no path to it, or a wrong-fit tool: cut or pause honestly. Being the person who says cut when cut is the truth is worth more than a forced renewal.
When the surrounding workflow needs attention
DORA's most durable finding is that AI acts as an amplifier. It magnifies the strengths of a healthy engineering org and the dysfunction of a struggling one. Faster coding does not automatically translate into more reliable delivery. Review, testing, deployment governance, and decision latency still affect the outcome.
When AI increases the supply of code faster than your org increases its capacity to review, test, and maintain it, the constraint simply relocates to the review queue. McKinsey's research points the same direction: fundamental workflow redesign, not tool choice, is the change most strongly correlated with real business impact, and only a small minority of firms have done it. When delivery is flat and review time rises, inspect the surrounding workflow as well as tool fit. These associations do not identify the cause for your team.
If the data is missing, state what remains unknown and start a comparable baseline. Weigh a limited pilot or temporary renewal against contract deadlines, switching costs, and the risk of disruption while you gather evidence.
Sources
Original research and publisher reports, with study limits and commercial sources identified.
- METR: experienced developers and early-2025 AI tools
Randomized trial; the 19% slowdown and perceived 20% speedup refer to this sample and period.
- METR: February 2026 productivity update
Follow-up with results consistent with speedups, but selection and timing biases prevent a reliable current estimate.
- DORA: 2025 State of AI-assisted Software Development
Google-led observational research; throughput and delivery-quality findings are associations.
- Faros AI: AI productivity research in 2025
Measurement-vendor telemetry from 10,000 developers; reports 91% longer review time.
- Faros AI: AI Acceleration Whiplash
Measurement-vendor telemetry from 22,000 developers; reports roughly five times the median review time.
- McKinsey: how organizations are rewiring to capture value
Survey associations with self-reported EBIT impact; does not establish that workflow redesign causes gains.
See where you stand.
Use the free self-check to identify which adoption, delivery, and review metrics you can already measure, and which ones need a baseline.
Questions leaders ask
What is the fastest way to save money at AI coding seat renewal?
Compare licensed seats with activity, then confirm each apparent exception with its owner. Leave and seasonal or planned work can explain inactivity. Review contract terms before estimating savings or deciding to remove, reallocate, or keep a seat.
Aren't my engineers saying they feel faster good enough to renew?
Not by itself. In METR's early-2025 sample of experienced open-source developers, perceived speed and measured completion time diverged. Its later update could not reliably estimate the current effect. Pair your team's feedback with delivery, quality, and cost data rather than assuming that result applies to them.
Does this decision method depend on which AI coding tool we bought?
The framework applies across coding tools, but available telemetry, metric definitions, and billing differ. Combine usage information with delivery, quality, and costs, and adapt the comparison to your team and plan.
What does "fix the system" mean instead of cutting the tool?
It means testing whether review, integration, or delivery practices are limiting the benefit. Flat delivery alongside larger PRs or longer reviews is a reason to investigate, not proof of a particular cause. DORA and McKinsey point to the importance of organizational context; use your own evidence to decide whether to change the workflow, the tool, or both.
What if we don't have any of this data?
Record the missing information and start a comparable baseline for delivery, review, quality, and costs. If renewal comes first, make the uncertainty explicit and consider a limited commitment or pilot while you collect evidence.