Key Takeaways
- →The Evaluation Gap: when leadership cannot judge production readiness, it rationally rewards what it can judge, so demos beat delivery.
- →Credentials do not close the gap. The evidence on technical CEOs is mixed in both directions; what varies is whether technical truth can reach the top and win an argument.
- →The base rates are a table problem: 95% of GenAI pilots show no measurable P&L impact, and RAND's number-one root cause is stakeholders miscommunicating the problem, not the model.
- →Reorgs on repeat are the tell. A table that cannot read the work rearranges the workers, and the new structure inherits the same gap it was built to escape.
- →Four Monday-morning questions close the gap without code: who defines production-ready, who can call a demo fake, what got celebrated last quarter, and what happened to the last three pilots.
Builder.ai entered insolvency in May 2025, after it cut the sales figures it had shown investors and a lender seized $37 million from its accounts. The collapse followed years of allegations, first reported by the Wall Street Journal in 2019, that human engineers were doing work the company presented as AI. Everyone wants to talk about the founders. I want to talk about every table that pitch passed through, because each one approved, funded, or bought on the strength of a convincing demo. The question is not how the story was sustained for so long. It is why no room it entered contained someone who could test it.
I've written about why so few leaders know whether their AI works. This is the governance root beneath that question. AI theatre persists wherever no one at the decision-making table can tell a demo from a product.
The Wrong Diagnosis
The reflex fix arrives within minutes of any AI embarrassment: put technologists in charge. It sounds like rigor. The record disagrees, in both directions.
Ginni Rometty holds a degree in computer science and electrical engineering from Northwestern. Under her tenure from 2012 to 2020, IBM logged roughly 20-plus consecutive quarters of revenue decline while the heavily marketed Watson business kept producing demos that never became durable products. A deeply credentialed CEO presided over the canonical demo-over-delivery era. I offer that as pattern evidence, and only as pattern evidence, because the pattern is the point.
Pat Gelsinger was the lead architect of the Intel 486. Credentials in this industry do not run deeper. Intel's board still pushed him out, announced December 2, 2024 (retirement effective December 1), after a stock decline of roughly 52 percent for the year and a foundry strategy that had not delivered volume. Hold that scene. It carries the freshest idea in this piece, and I'll come back to it.
Now flip the coin. Brian Chesky took a fine arts degree in industrial design from RISD in 2004 and built Airbnb to a public debut whose first-day trading valued the company above $100 billion. Stewart Butterfield studied philosophy, a bachelor's and then an MPhil, and built Slack, which Salesforce bought for $27.7 billion in a deal that closed in July 2021.
The research is just as stubborn as the anecdotes. Ali Tamaseb's dataset of billion-dollar startups found roughly equal numbers of technical and non-technical unicorn CEOs, with similar success. Bhagat, Bolton and Subramanian, working across more than 14,500 CEO-years, found no significant link between a CEO's education and firm performance (studies of inventor and STEM CEOs, Islam and Zein in 2020 among them, do find innovation advantages). The credential evidence is mixed. Say it clean and let it go.
Yet if credentials don't predict the outcome, something else varies. What varies is whether technical truth can reach the top of the organization and win an argument.
The Evaluation Gap
Here is the definition I'd ask you to carry into your next board meeting. The Evaluation Gap is what opens when leadership cannot judge production readiness, so it rationally rewards what it can judge.
Walk the incentive logic slowly, because it has to survive being repeated in a meeting. Organizations optimize for what their leadership can verify. Any executive can verify press coverage. Any executive can verify demo polish and a launch event. Verifying uptime, evaluation results, and cost per served request takes fluency, or a trusted translator with real standing. Whichever signal the table can read is the signal the organization learns to produce.
I picked the word rationally on purpose. This mechanism runs without villains. Smart people respond to the signals that actually reach them, organizations repeat whatever earned applause last quarter, and the loop closes with everyone behaving sensibly. That is also the exit ramp: a structural problem admits a structural fix.
Follow the chain downstream. Wrapper companies get funded because they optimize the legible layer. Pilots stall in purgatory because nobody at the table can define done. The gap even gets policed from outside when nobody polices it at the table: the SEC charged Delphia and Global Predictions with AI-washing in March 2024, then Presto Automation in January 2025, the first public company on that list. And when delivery fails without a legible cause, the failure gets read as a people problem.
A board of engineers is beside the point. The gap closes when technical truth has a seat, a voice, and the power to win an argument. The four questions at the end of this piece test exactly that.
The Gap, Measured
The base rates deserve one paragraph in one place. MIT's NANDA initiative reported in "The GenAI Divide: State of AI in Business 2025" that 95 percent of enterprise GenAI pilots show no measurable P&L impact. RAND found in 2024 that, by some estimates, more than 80 percent of AI projects fail, twice the rate of non-AI IT projects. S&P Global Market Intelligence, surveying more than 1,000 enterprises in 2025, found that 42 percent of companies had abandoned most of their AI initiatives, up from 17 percent the year before, and that an average of 46 percent of AI proofs of concept were scrapped before production.
Numbers like these usually get filed under doom. Read them through the mechanism instead. RAND's number-one root cause is stakeholders misunderstanding or miscommunicating what problem needs to be solved. That is not a model failure. That is a table failure.
Boards themselves have now been measured. MIT Sloan Management Review reported in December 2025 that only 26 percent of 2,800 large public companies have AI-savvy boards, and that those firms post return on equity 10.9 percentage points above their industry average while the rest sit 3.8 points below it. Their finding is the correlation between board savviness and performance. Mine is the mechanism underneath it, what a table rewards when it cannot evaluate, which also keeps this out of Founder Mode territory: the variable is evaluation capacity, wherever the founder happens to sit.
Money makes the cleanest stress test, because if capital alone closed the gap, state-scale funding would have closed it. The sovereign-AI class had the capital. Mistral stands as its lone commercial success, past $400 million in annual recurring revenue as of February 2026. Germany's Aleph Alpha abandoned frontier models and merged into Cohere in April 2026, with shareholders receiving roughly 10 percent of the combined entity. India's Krutrim pulled its consumer app in April 2026 after roughly 200 layoffs, and Sarvam-M managed about 334 Hugging Face downloads in its first two days. The EU's Teuken-7B has logged roughly 100,000 all-time downloads of its main commercial variant on Hugging Face (counter read July 31, 2026), and I could find no publicly announced enterprise deployment. Naver was eliminated from Korea's own sovereign-AI program in January 2026. One success against a field of stalls. Capital does not buy evaluation capacity.
The same spread shows up at market scale. A McKinsey survey covered by The National in November 2025 found that 84 percent of GCC organizations have adopted AI in at least one function while only 11 percent generate measurable financial returns. Markets where adoption gets announced faster than it gets measured show how far activity can outrun return.
The Tell: Reorgs on Repeat
When an organization cannot evaluate the work, it evaluates the people. Delivery failure arrives with no legible cause, and the only lever the table knows how to pull is the org chart.
I've seen the same pattern on three continents. A stalled AI program is followed by a reshuffle. Then a renaming. Then a new operating structure with a fresh acronym and the same people two boxes over. The new structure inherits the exact evaluation gap it was built to escape, so some quarters later the org chart gets redrawn again. One reshuffle is a Tuesday. The cadence is the tell.
Reorganizations carry their own sobering literature. HBR and McKinsey work documents that most reorganizations fail to deliver their intended value on the timeline planned. What I have not found anywhere is the connection between reorg cadence and AI delivery failure, so let me make it as a bounded observation: a reorg every few quarters around the same stalled AI program is the visible symptom of the Evaluation Gap. The table cannot read the work, so it rearranges the workers.
Which returns us to Gelsinger. When Intel's foundry strategy had not delivered volume, the board's move was to change the person. I offer that as my reading of the public record, nothing more. Changing people is the move available to a board that cannot adjudicate the technical dispute itself.
Freezing the org chart fixes nothing either. Make the work evaluable, and the org chart stops being the only instrument the table knows how to play.
Four Questions for Monday Morning
None of these four requires a technical vocabulary. Each tests whether technical truth can reach this table and win, the thing the credentials debate was never about.
1. "What does production-ready mean for this system, and who at this table can define it?" It forces a definition of done before more money moves. Silence is itself the finding.
2. "Who is empowered to tell us this demo isn't real, and when did they last say so?" An empowered truth-teller who has never once used the power is a decoration.
3. "What did we celebrate last quarter: an unveiling or uptime?" This one audits the incentive structure in a single glance. Whatever gets applause gets repeated.
4. "What happened to our last three pilots after the announcement?" This converts anecdote into a base rate the board owns. For calibration, across more than 1,000 enterprises an average of 46 percent of AI proofs of concept were scrapped before production, so a board that cannot answer has plenty of company. It just has no protection.
A board that asks these four questions, consistently, changes what the organization optimizes for. Not one member has to write a line of code to do it.
The Other Failure Mode
Fairness demands the mirror image, so here it is. Engineers fail in the opposite direction: we build over-engineered platforms nobody asked for, fall in love with the architecture, and act surprised when the business shrugs. I write that from inside the tribe. The mixed evidence on credentials cuts both ways, and it should, because what decides these outcomes is who wins the argument when the demo and the truth disagree.
That takes three things. Fluency at the table. Truth-tellers with real power. Incentives that pay for deployment instead of unveiling.
The last piece in this series asked whether you know if your AI works. This one ends a level deeper, where governance actually lives. Could anyone at your table prove it, either way?
Question four has a base rate; your table should have one too. Run the AI Readiness Canvas: 15 questions, your score in 10 minutes, and a read on whether your organization can tell a demo from a product.
Get Weekly Thinking
Join 2,500+ AI leaders who start their week with original insights.

Senior AI strategist helping leaders make AI real across four continents. Forbes Technology Council member, IEEE Senior Member.
Ajay's views, from 15 years in the field. Not legal or compliance advice. See full disclaimers →
Published by AI Exponent LLC