← All servicesFor engineering teams

If your Claude Code rollout is not showing a return.

Three fixed-scope engagements led by an AI Coding ROI Readout. Published pricing, honest measurement, artifacts you keep. Built for engineering orgs of roughly 20 to 500 engineers. If you are the individual engineer, the 1:1 coaching on the services page is probably closer to what you need.

Ex-Amazon · Ex-The New York Times · Prices published below

The engagements

Start with the Readout.

One fix-diagnostic front door, then the deeper engagement it points to, and a standalone workshop for the house style that makes it stick. Full detail on each engagement’s own page.

Start here
~2 weeks · mostly async

AI Coding ROI Readout

$4,500
3 tiers · Startup $2,500 to Extended scoped

A 2-week diagnostic that finds exactly where your system is leaking the AI spend (review bandwidth, workflow, guardrails), and the on-ramp to installing the fix. You keep a board-ready deck and a keep, renegotiate, fix, or cut recommendation.

  • Adoption and throughput data pull
  • Benchmark vs. independent medians
  • Spend versus value, by team
3-week engagementOnboarding founding clients

AI-Native Delivery Accelerator

$7,500
3 tiers · Startup $5,000 to Enterprise scoped

The deeper engagement the Readout points to: install the operating-model fix and prove it on one team and one repo, with a before-and-after number.

  • Week 1 · Diagnose
  • Midpoint · Half-day readout
  • Weeks 2 to 3 · Execution support
90 minutes, live · remote or onsite

CLAUDE.md Workshop

$3,000

A live 90-minute working session that encodes your team's standards into your real repositories, so coding agents follow the way you already ship: context files, review habits, and workflows tuned to your codebase.

  • A live 90-minute working session, on your repos
  • Your CLAUDE.md files, written live
  • Review habits your team will keep

First engagements open soon · Founding-client pricing for the first two teams

Adoption was never the problem.

I was on the team that rolled out Claude Code to The New York Times Engineering. Your engineers are using the tools. The reason you cannot point to a return is that the bottleneck moved downstream, and the system never moved with it.

  • 01

    More code, same review

    AI pushes more code through a review process built for a slower kind of output. The queue backs up, and cycle time stays flat or gets worse.

  • 02

    The bottleneck moved downstream

    The constraint was never writing code. It is now review, risk-tiering, and measurement, and none of that was redesigned for machine-speed output.

  • 03

    Nobody can point to the return

    Throughput is up and the spend is real, but no one can tell the CFO what the seats actually bought. What you cannot measure, you cannot defend.

The fix is not more usage. It is redesigning the system of work around the code: review, standards, measurement, and guardrails your teams actually adopt. That is what these engagements install.

The method

A six-layer diagnostic, not a tool demo.

AI amplifies the system of work you already have. Point a coding agent at a team whose review, measurement, and workflow were built around humans reading every line, and it magnifies the dysfunction instead of fixing it. The diagnostic runs six layers, starting with the one most teams skip: knowing what the tool actually changed.

  1. L0

    Honest measurement

    Most orgs cannot even state what the tool changed. Instrument real delivery and developer-productivity outcomes before scaling anything, not license and acceptance rates.

  2. L1

    Redesign the workflow

    AI bolted onto your old lifecycle wastes it. Move its value upstream into specs, design, and tests, and keep batches small and reviewable. The highest-ROI lever.

  3. L2

    Risk-tier the review

    Not every change is equally risky. Auto-review low-stakes diffs and concentrate senior judgment, with CODEOWNERS, on sensitive paths like authentication and deploy config.

  4. L3

    Architecture and context

    A starved agent cannot see a tightly-coupled codebase. Codify architectural intent and give the repo real edges the agent cannot cross.

  5. L4

    Platform and guardrails

    Quality slips in at scale. Enforce tests, security scans, and policy gates at the merge boundary, not just locally, and trace AI-written code to production.

  6. L5

    Culture and enablement

    Adoption stalls despite paid seats. A clear AI policy, hands-on enablement, and measuring at the team level, not rewarding raw diffs per engineer.

Independent data

The numbers vendors won’t show you.

We anchor to independent and published research: METR, DORA, McKinsey, Meta RADAR, and Carnegie Mellon. None of it supports the vendor 10x pitch.

The pitch is 10x. Developers feel about 20% faster. Independent research from METR and DORA finds real team-level gains are modest and uneven, and org-level output barely moves. That gap is where your spend leaks, and measurement is how you find it.

METR· independent
felt +20%, measured −19%

In METR's controlled 2025 trial, experienced developers believed AI made them about 20% faster while it actually made them 19% slower. Perceived productivity is an unreliable instrument, which is why we measure outcomes, not how fast your team feels.

Meta RADAR· published research
1/3 the revert rate

Meta's risk-aware review auto-approved low-risk diffs and routed high-risk ones to humans. Those governed diffs saw one-third the revert rate and one-fiftieth the incident rate. Governance must evolve, and human judgment belongs on the high-stakes paths.

DORA· independent
amplifier, not accelerator

DORA reframes AI as an amplifier that magnifies a healthy org's strengths and a struggling one's dysfunction. Even as throughput turns positive, delivery stability still tends to degrade, so fix the system before scaling the tool.

McKinsey· independent
workflow redesign

McKinsey's State of AI 2025 found that fundamental workflow redesign is the organizational change most correlated with AI's bottom-line impact. That is why we lead with workflow and measurement, not tooling.

Carnegie Mellon· independent academic
~41% higher complexity

Carnegie Mellon researchers studied 807 repositories and found AI coding tool adoption associated with higher code complexity. More code can carry a downstream cost, so the system around the code has to change too.

Not a dashboard

A platform can’t tell you to cancel it. I can.

Measurement platforms are strong annual tools you install and staff, at roughly $50,000 to $350,000 a year. The Readout is a different tool for a different job: a one-time decision, not a subscription you renew to find out whether a subscription is worth it.

  • 01

    A fixed price, a one-time decision

    Not an annual subscription to find out whether your annual subscriptions are worth it. You pay once, you get an answer, you act on it.

  • 02

    No product to renew, so the answer can be cut

    If the honest read is that a tool is not paying off, I will tell you to cut it. A platform whose revenue is the measurement never will. That is what real neutrality looks like.

  • 03

    A decision, not another dashboard

    You leave with a board-ready narrative and a keep, renegotiate, fix, or cut recommendation, not one more dashboard your leadership still has to interpret.

  • 04

    Zero procurement drag

    Read-only, in your environment, code never leaves your infrastructure, done in about two weeks. No security review, no data piped to a third party.

Process

How the team engagements work.

  1. 01

    30-minute discovery call

    Free, no pitch. Your setup: seats, tools, team size, who owns the rollout, and where it is stalling. You leave with something useful either way.

  2. 02

    Pick the engagement

    I propose the smallest engagement that fixes the problem and name the price on the call. No fit, no proposal: you get a one-page write-up of what I saw instead.

  3. 03

    Fixed scope, fixed dates

    Every engagement has an end date set at kickoff. Async support runs in a shared Slack channel with next-business-day responses.

  4. 04

    You keep the artifacts

    Decks, recommendations, templates, recordings. Everything is built to run without me after the engagement ends.

Straight answers

Questions teams ask.

$4,500 is a lot for a report.

It is the diagnostic on-ramp to the fix: where your AI-coding spend is leaking, what system change would recover it, and whether to keep, renegotiate, fix, or cut. The full fee is credited toward the Accelerator if you decide to go deeper within 30 days. If the price still gives pause, the Startup tier is $2,500.

We already measure this with our own tooling.

Good, most teams do not, so you are ahead of the field. The Readout then benchmarks what you already have against independent research from DORA and METR and pinpoints where the value is leaking, which is usually downstream of the individual coder. If you track license and acceptance rates but not delivery outcomes, that gap is exactly the wedge.

Why you instead of a big consultancy?

You get a practitioner who ships under this model daily, not a slide deck from someone who has never merged an AI-written pull request. At The New York Times I was on the AI Champions team that rolled Claude Code out across the roughly 500-engineer org, delivered three presentations including a CLAUDE.md talk at the internal EngX conference, and was the second-most-prolific contributor to the internal Claude Code plugin marketplace. The work is tool-agnostic and every number is anchored to independent data.

Can we just do this internally?

Your internal champion already exists, and adoption still stalled, which is precisely why there is pain. An outside seasoned read and a sanctioned, org-wide standard are things an insider cannot self-grant, because no single engineer can authorize the whole team to change how it works. I make that call easier to sanction, then hand it back to your team to run.

Is this Claude Code only?

No. Claude Code is the tool I know deepest, so it is the natural wedge, but the methodology is tool-agnostic and built on durable principles rather than any one model or agent. The bottlenecks it fixes, review bandwidth, missing baselines, and a workflow built around humans reading every line, are the same whether you run Claude Code, Copilot, Cursor, or whatever wins next.

What results should we actually expect?

Honesty is the whole point, so there are no 10x promises. Independent research from METR and DORA sets the frame: gains are modest and uneven, and the larger wins come from clearing the downstream bottleneck, not from speeding up the coder. The Accelerator proves it on your own code with a one-team pilot and a before-and-after delta on review time, throughput, and revert rate, and that number becomes your attributable case study.

How do you handle code access and security?

The front-door Readout is deliberately low-access: dashboards and interviews, not a repository deep-dive. For anything that touches code, I work read-only inside your environment, time-boxed, so code never leaves your infrastructure, under a mutual NDA. I state delivery practices plainly and never claim a certification I do not hold.

What size team is this built for?

Engineering orgs of roughly 20 to 500 engineers, with the mid-market sweet spot around 20 to 200, where a coding-agent rollout is already live but the system of work around it was built for humans reading every line. The front door is sized to approve on a team budget, so the buyer is usually an engineering manager, director, or Head of DevEx, not a procurement committee. Mid-market software and media teams usually need enablement and measurement most; regulated enterprises need the risk-tiering most. If you are under a dozen engineers on greenfield code, you likely do not need this yet.

Tell me where the rollout is stalling.

Seats, tools, team size, who owns the rollout, and where it is stuck. The discovery call is 30 minutes and free. You leave with a read on the shape, whether or not we end up working together.

alextongme@gmail.com

I reply within one business day