The Steelman
Take both numbers at face value for a moment, because the case for doing so is stronger than the case against.
Anthropic did not estimate. It sampled a fifth of the staff in every department, every week of July, had an agent reconstruct what each person actually worked on from Slack and internal documents, and built a taxonomy of 542 nodes before rating anything. That is more care than most firms apply to counting their own headcount. Faros did not run a survey either. It read two years of telemetry from 22,000 developers.
Set against that, objecting that a number is imprecise sounds like pedantry, and worse, convenient. Every executive who would rather not act on AI has a reason the data is not good enough yet, and "your methodology has error bars" is the most respectable one. Past a point, demanding a better measurement is a way of avoiding the one you have.
So here is what both organisations did that the coverage of them did not.
Anthropic published the number and its disagreement rate in the same post. Claude leads 26% of its AI R&D work, up from under 1% in February. Then, further down the same page: when Anthropic asked the staff who own that work to rate it themselves, those people agreed with each other on the exact level 35% of the time. The model judge agreed with humans 59% of the time. Humans agreed with humans 35%.
Read that ordering again. The machine was the more consistent rater. Not the more correct one, the more consistent one, which is a different and worse problem.
Faros did the same. Code churn up 861% under high AI adoption, a figure travelling since April as evidence of what AI rework costs. The report declines that reading. It lists three candidate explanations, rework, large-scale refactoring previously too expensive to staff, and engineers finally replacing code they never liked, then says the metric cannot resolve the ambiguity.
Neither caveat was caught by a critic. Both were volunteered, in the document, beside the headline. I read both end to end before writing this, and neither is buried. In both cases the number travelled and the sentence underneath it stayed home.
Anthropic rates every task category on six levels. The point estimate sits on "leads". The band is the honest reading: model and human ratings landed within one level 97% of the time.
The Decoder
An automation level is a judgement, not a reading off an instrument. The six rungs run from no AI involvement, through minimal, assists, collaborates and leads, to autonomous. Nothing in that ladder is counted. Someone looks at a week of work and decides which rung it sits on.
Which is why the 35% matters more than the 26%. It is not telling you the raters were careless. It is telling you the category is genuinely hard, that reasonable people looking at identical evidence place it differently, and that the honest output is a band rather than a point. Anthropic's own within-one-level figure, 97%, is the band. That is the number a serious reader should quote.
A measurement that cannot tell you how much it disagrees with itself is not a measurement. It is a slogan with a decimal point.
When a team reported a large productivity gain from an agentic coding deployment, I asked how they were counting it, and every leader in the room turned out to have built their own scale around their own mandate, so we could not agree on what one unit of work was.
The unit of work is the argument.
The index and its limitations section are both worth twenty minutes: read what Anthropic published about how it measured, not the write-ups of it.
The Playbook
What Faros published: one figure, three live explanations, none selected.
Ask for the denominator before the percentage. When someone reports that AI now does some share of the work, ask what one unit of that work is and who decided. If the answer runs longer than a sentence, the number is not ready to repeat upward.
Require a disagreement rate on any judged metric. Anything scored by a rubric rather than counted by a system needs two raters on a sample and a published agreement figure. Take thirty items, score them independently, report how often they matched. Below half, you have a definition problem, not a performance result.
Report the band, not the point. Anthropic's 97% within-one-level is a better sentence than its 26%, because it survives contact with a sceptic.
Price what the demo does not show. In that room nobody had netted out what the productivity number hides. Rework when output comes back wrong, which is the real cost of AI. The supervision time someone absorbs quietly. And ownership: who answers by name when the agent is wrong. A gain netting none of these out is a timing difference.
Separate what shipped from what survived. Throughput is easy to instrument and tells you little on its own. A rework rate means paying for provenance, at the line or the ticket, which Faros says plainly its own dataset could not do. Budget for that or stop quoting rework.
The Signal
Watch which half of each disclosure gets quoted over the next month. Both organisations behaved well. They shipped the uncertainty with the finding, which is the standard everyone claims to want. The market rewarded them by keeping the headline and dropping the caveat.
I think the question worth carrying into Monday is not how much work AI is doing, but whether anyone in the room can say how they counted. If candour is reliably punished, the next lab or vendor with a number to publish will notice. Consistency was the machine's advantage in Anthropic's own data. Candour should not turn out to be a disadvantage in ours.
P.S. To run this on your own numbers, naming your units of work is where the AI Strategy Canvas starts.
The Correction
As reported: Coverage of the Faros AI 2026 engineering report has treated its 861% rise in code churn under high AI adoption as a measurement of AI rework, and its 54% rise in bugs per developer, against 9% in the prior report, as a year-on-year trend that is steepening.
What the source says: The report declines both readings itself. On churn it gives three candidate explanations, rework, newly feasible large-scale refactoring, and engineers replacing code they were never satisfied with, and states the metric "cannot resolve the ambiguity"; a footnote adds that separating the two requires line-level provenance data the dataset does not hold. On the year-on-year comparison it states the two studies are "independent cross-sections of the industry, not a longitudinal panel". Faros AI, The Acceleration Whiplash, April 2026
The Framework
This is exactly why I built the The Agentic Work-Unit Ledger.
The Ledger opens on the work unit and closes on the rework rate, and this issue is what both lines look like when someone tries to satisfy them in public. Naming the unit costs you apparent certainty, and that cost is the point.
Your Agent Isn't a Feature →On My Radar
- Faros also reports 31.3% more pull requests merging with no review at all, human or agentic, which it calls the most urgent finding in its section.
- Anthropic is explicit that Claude is not operating fully autonomously for any measured subset of its AI R&D work, a line largely absent from the coverage.
- The Ledger argues an agent is a hire rather than a feature, which is why the unit of work has to be named before it can be priced.
I would rather carry a band I can defend than a point I cannot. If your AI numbers do not travel with one yet, that is this fortnight’s work.
Know someone who needs this? Forward it to the engineering leader who just reported an AI productivity number upward.
Ajay's views, from 15 years in the field. Not legal or compliance advice. See full disclaimers →