Key Takeaways
- →Agentic AI does autonomous work, so measure it like labor, not software: cost per unit of finished work, not cost per seat.
- →The Agentic Work-Unit Ledger has five lines: three you count (work unit, loaded cost per unit, rework rate) and two you answer in words (redeploy-or-reduce, the named human).
- →Rework rate is the quality tax vendors hide. Close the outcome partition to 100% and watch the deflection bucket you are not shown.
- →Loaded cost per unit is a curve, not a point: report the trailing average and its worst month, and sample supervision instead of guessing it.
- →Used cynically the ledger is a firing spreadsheet; used honestly it forces a redeploy-or-reduce disclosure with a named human, never a laundered "yield."
Last year I argued that most leaders count the wrong things about AI ROI. I was right. I was also already out of date. Agentic AI has changed which things.
That earlier piece ("Measuring AI ROI: Why Most Count the Wrong Things") walked through the machinery for evaluating a deployed AI feature: an enhanced cost-benefit analysis where the "often missed" column routinely carries 40 to 60 percent of the real value, a balanced scorecard with a process quadrant, predictive modeling that hands boards a range instead of a false-precision point, and the J-curve that talks nervous executives out of killing a winner at month six. I stand by all of it for what it was built to measure. A tool.
An agent is not a tool. You do not deploy it and amortize it across seats. It does work. It resolves the ticket, reconciles the account, reviews the contract, matches the invoice, and it does that work without a human holding the wheel for every turn. The moment software stops assisting a person and starts producing finished output on its own, you have quietly moved a line item from your technology budget to your labor budget. And you are still measuring it with the technology playbook.
So here is the reframe the whole piece rests on. An agent is a hire, not a feature you deploy. You measure a hire with a ledger of the work, not a payback curve.
The category error
Feature ROI answers one question well: was the money we spent on this capability worth the value it returned? Buy once, amortize, count the seats, watch the J-curve, book the payback. That model isn't wrong, just aimed at the wrong object.
A worker generates a different set of questions. How much does one finished unit of their work cost, fully loaded? How often does it come back wrong? What happens to the person whose work they absorbed? None of those are capability questions. They are labor questions, and the honest answers move around week to week in ways a purchase price never does.
Here's the part I want to be precise about, because a careful reader of my earlier framework will call me on it otherwise. The old model does not ignore supervision and rework. It has an "often missed" column and a process quadrant, and a diligent analyst can file both costs there. My claim is narrower and, I think, more useful. The old model treats supervision and rework as afterthoughts. A footnote. A sub-quadrant you get to if you have time. For a feature that runs on rails, that placement is fine. For a worker that acts on its own judgment, supervision and rework aren't a footnote. They are the two biggest lines on the page. The Agentic Work-Unit Ledger promotes them from afterthought to load-bearing. That promotion is the entire sequel.
The Ledger, line by line
Five lines. Each one carries an operating rule that turns it from a nice idea into something you can actually run on Monday.
Line 1, the work unit. Name the atomic finished output the agent produces. A resolved support ticket. A reconciled account. A reviewed contract. If you can't name the unit, you can't measure it, and you should stop here until you can. Then hold the definition still. This is the fixed-unit rule, and it matters more than it sounds. As an agent gets good, it silently absorbs the easy cases, and only the hard ones keep reaching it. Your cost per unit will rise. Read that as success, not decay: the machine ate the trivial work and left you the residue that was always going to be expensive. When scope expands like that, you don't get to keep comparing against the old number. You start a new ledger. Comparing this quarter's hard-case cost to last quarter's easy-case cost is how you talk yourself into killing something that's working.
Line 2, loaded cost per work unit. Tokens, plus orchestration, plus tooling and integrations, plus human supervision, divided by units of finished work. Tokens are the smallest part of this bill. Supervision is the line software ROI never had to carry, because software did not need a human watching in case it improvised.
Two things make this line honest. First, it is a curve, not a point. Report the trailing average and its worst month, side by side. Supervision cost is U-shaped and incident-driven: high while you onboard, low through the honeymoon, then it spikes the first time the agent fails silently at scale and the org over-corrects into review-everything mode. A single average smears out the tail that is the whole story. Second, you do not need a timesheet code to measure it. Sample it. For a week or two, one reviewer logs supervision minutes per unit, you annualize the ratio, and you re-sample after any major model or scope change. That defeats the objection I hear most, which is that line 2 is uninstrumentable. It is not. It is just un-sampled.
This is also where the capability numbers belong, read correctly. METR's updated time-horizon work (8 May 2026) shows frontier agents completing tasks up to roughly sixteen hours of expert-equivalent effort. The number people quote and stop. The number that matters is the reliability attached to it: that ceiling holds at 50 percent reliability on a software-task suite. Half the time, it needs a human. Read as a capability flex, sixteen hours is a headline. Read honestly, it is a supervision-cost argument, and the honesty is the point.
Line 3, rework rate, the quality tax. The share of units escalated, corrected, redone, or bounced back. This is the line that separates a real agent from an expensive demo, and it's the one vendors are best at hiding.
Two rules keep it clean. First, close the partition. Every unit lands in exactly one bucket, and the buckets sum to a hundred percent: resolved clean, reworked, escalated, or deflected-and-abandoned. Deflected and abandoned units count against the agent unless someone independently confirms they were actually resolved. An agent that "handles" a ticket by getting the customer to give up is not doing the work. It is hiding it. Second, watch the bucket you aren't shown. If your vendor reports a rework rate but not a deflection rate, the tax is hiding in the number they didn't put on the slide.
Line 4, redeploy or reduce. You don't compute this one. You answer it in words, with evidence, in the annual operating review, and it is the honest core of the whole instrument.
When an agent absorbs a body of work, you did one of two things with the person who used to do it. You redeployed them to higher-value work, or you reduced the role. The Ledger's one demand here is that you say which. Booking a headcount cut as "reinvestment yield" or "freed capacity" is laundering, and I'm naming it as laundering on purpose. There's a comfortable version of this framework where absorbed work becomes abstract "slack" that quietly reappears as un-captured ROI. I used to reach for that framing. I don't anymore. It's the sentence that lets an organization avoid looking at what it actually did.
Line 5, the named human. The Ledger measures tokens and orchestration and supervision as costs. It must also name who carries the rework tax, because that's a real person with a real role, and it must state what happened to the worker whose output the agent absorbed: retrain, redeploy, or reduce. Not folded into a yield number. Named.
I'll say the uncomfortable part before someone says it for me. Used cynically, this ledger is a firing spreadsheet. Line 1 names the work, line 2 prices the human out of it, and line 5 tells you which humans. I know exactly how it reads in the wrong hands, and I'm not going to pretend the risk away. So here is the guardrail that separates the two uses. An honest ledger states its workforce disposition out loud, in the operating review, where the board and the affected function can both see it, and it treats "reduce" as a decision with a name on it, not a rounding error inside a productivity metric. A firing spreadsheet hides that line. The Ledger requires it. You can't stop a hostile crop of this article. You can make the crop require dishonesty to land, and that's the only real defense any framework has.
Where this belongs, and where it doesn't
A fair CFO will push here, so let me meet the push directly. Not all AI spend is agentic. When AI assists a person who still owns the output, the draft they'll rewrite, the code they'll review, you're in augmentation, and augmentation stays with the old feature-ROI model. Nothing here replaces that. Treat the boundary as a trajectory, not a wall I'm building to protect a new framework. Agentic work is the fastest-growing and worst-measured slice of AI spend, and it's where the next wave of blind money is going. KPMG's Q2 2026 global survey of more than two thousand senior leaders found only 7 percent had established clear ROI, 42 percent had merely partial visibility into their AI spend, and 23 percent couldn't see usage-based costs at all. The organizations with real cost visibility were five times likelier to show ROI. This instrument is for the money you're about to spend on autonomous work, not the money you already track.
And I'll concede the lineage before anyone accuses me of dressing up an old idea. Yes, this is activity-based costing pointed at a labor unit that didn't exist eighteen months ago. The bones are ABC. What's new lives in three lines that no BPO invoice and no software ROI model has ever carried: the rework tax as a first-class number (line 3), the honest redeploy-or-reduce disclosure (line 4), and the named human (line 5). Take those three away and I'm just billing you for output. Those three are the reason I bothered.
There's a real gap here that no current measure fills. Anthropic's Economic Index (June 2026) shows how deeply AI now reaches into daily work. Most users report gains in the speed, scope, and quality of what they produce, and more than a third expect AI to handle most of their tasks within a year. What it explicitly does not measure is what those workers did with the time that got freed. That gap, the space between "the task was absorbed" and "here's what happened to the person," is exactly the space lines 4 and 5 exist to fill. Gartner has separately described enterprise AI agents entering a phase of more complex pricing and ROI dynamics as billing shifts from seats toward consumption. Seats were a feature signal. Consumption is a labor signal. The measurement has to move with the money.
One worked example, honestly
Picture a support-ticket agent. Hypothetical, no real company. Tokens are almost free, a few cents a ticket, and the demo is dazzling. Then you run the Ledger. Rework rate comes in at 40 percent once you count the tickets that bounced back and the ones the bot quietly deflected into a dead-end FAQ. Loaded cost per resolved ticket, with the reviewer who now reads every escalation, lands well above the cheap-token story. Three months in, scope has drifted: the agent absorbed all the password resets and only angry billing disputes reach it now, so cost per unit climbs, and you correctly read that as success and open a new ledger rather than panicking. And the reckoning: two support roles' worth of work got absorbed. Did those two people move up into complex-case handling, or did they leave? The Ledger doesn't let you skip the question. Feature ROI never asked it.
The one board question
Stop asking what the agent cost. Ask what a unit of its work costs, fully loaded. Ask how often you redo that work. And ask, honestly, what happened to the person whose work it took. If you're carrying autonomous agents on a payback curve built for software, you are measuring a hire like a purchase, and the gap between those two things is exactly where the money and the ethics both go missing.
If you want to run this ledger against your own agent portfolio, that is the work I do in agentic transformation engagements.
Get Weekly Thinking
Join 2,500+ AI leaders who start their week with original insights.

Senior AI strategist helping leaders make AI real across four continents. Forbes Technology Council member, IEEE Senior Member.
Ajay's views, from 15 years in the field. Not legal or compliance advice. See full disclaimers →
Published by AI Exponent LLC