Doing Less With More
Part 4 of Retcon Reckoning (The Great AI Replacement)
The gap between what the invoice said and what the board slide promised was not, in the end, a measurement problem. It was the measurement itself.
The corporate math was presented, in every instance, as self-evident. Headcount is a cost. AI reduces the need for headcount. Reducing headcount reduces cost. The freed cost funds AI investment. The AI investment further reduces headcount. The model is recursive. The flywheel spins. The slide is clean. The coffee is ok.
Gartner surveyed 350 large companies actively deploying AI, by any reasonable standard the organizations most positioned to be getting this right, and found that 80 percent had reduced their workforce as a ‘direct result’ of their AI investments. The finding that followed belied what assumptions those announcements would have otherwise invited: the layoffs had zero statistical correlation with improved ROI. Companies that cut the most staff showed returns nearly identical to companies that cut the least.
One mechanism the correlation data cannot explain is knowledge debt. Not the kind that appears on a balance sheet, bc it doesn’t appear anywhere a dashboard monitors or a quarterly review surfaces but is the accumulated understanding of a specific system, built by specific people over specific years of decisions that were made for reasons that made sense at the time and were never written down because the person who made them was always available to explain them, until they weren’t. The lag between cutting the humans and discovering what the humans were doing is long enough that the decision looks rational at execution. The first quarterly review looks fine. The second looks fine. The third surfaces something that cannot be explained without reference to a person who left eight months ago, and by then the documentation they were supposed to produce before leaving has been discovered to be incomplete in a particularly crucial way, bc the knowledge that matters most is hardest to articulate, which is why it was never articulated, and why it left with them. Which is why we aim organisationally to eliminate irreplaceability.
What the Gartner data also cannot capture is the response from the other side of the transaction. One study found that 29 percent of employees admit to actively sabotaging their company’s AI strategy, with the number climbing to 44 percent among GenZ, whose resistance is deliberate and tactical: feeding corrupted data into models, routing around corporate control planes with unauthorized tools, intentionally generating low-quality output to demonstrate system fragility, all while 60 percent of executives plan to lay off employees who won’t adopt AI, 77 percent will bar non-adopters from promotion, and 92 percent openly admit they are cultivating an AI elite. Leadership has spent two years insisting this is a productivity initiative. The workforce isn’t so sure, and neither is the data.
Moreover, the data does not and cannot capture what is happening to the work itself. The model does not argue when you accept its suggestion for another layer of indirection. It does not push back when the product requirement is vague and the prompt is lazier still. The model is the world’s most expensive yes-man, and organizations have discovered they like being told yes. The result is not acceleration toward excellence but acceleration toward sufficient because sufficient ships, sufficient meets the KPI, and sufficient can be presented to the board as evidence that the AI investment is bearing fruit, even as the long-term carrying cost of the sufficient codebase (quietly) metastasizes in production.
First The Good News
Q1 2026 earnings calls established is that something is happening.
What they did not establish is what.
Zuckerberg described engineers building in a week what once required dozens and months. IBM disclosed 45 percent productivity gains across its developer workforce and $4.5 billion in internal cost savings since 2023. These are quantitative claims in the technical sense, containing numbers, stated with executive confidence on a live call to institutional investors. They are not quantitative claims in the empirical sense, because none were independently verified, and because the people making them had a structurally motivated interest in demonstrating that the AI investment was generating returns.
But Meta did post the most lucrative opening quarter in its corporate history ($56.31 billion in revenue, $26.8 billion in net income) and three weeks later announced 8,000 layoffs. At the internal town hall, Zuckerberg was direct: getting everyone to use AI tools and doing the work more efficiently is not the thing that’s driving layoffs. The CEO who had just spent an earnings call crediting AI-enabled productivity gains, now telling his own employees that the productivity gains were not why anyone is losing their job. Cisco ran the same play this quarter. Cloudflare followed suit with an apologetic memo described as “the first true AI-layoff manifesto.”
So, AI is enabling companies to do more with less (according to the earnings calls) but also AI is not the reason for all the pink slips (according to the layoff announcements.) And then there’s the CEPR survey of 5,000 firms in the same period found that around 80 percent reported no AI-driven productivity gains at all. The Atlanta Fed found measurable gains concentrated in high-skilled tasks, with perceived gains exceeding measured revenue gains, a gap the researchers called the productivity paradox, which might less diplomatically be described as the distance between what gets said on earnings calls and what shows up in the books. McKinsey found only 39 percent of companies reporting any current EBIT impact from AI. But 39 percent is not nothing.
The actual data resolves into a measurement fog dense enough that the same quarter can produce IBM’s $4.5 billion in savings and a survey of 5,000 firms reporting nothing measureable. Both can be accurate while neither indicate what is actually happening, because the proxies measure things adjacent to the outcome rather than the outcome itself ; the metric has decoupled from what it was introduced to track, and the decoupling is invisible from inside the metric. When Uber rolled out Claude Code to its 5,000-person engineering organization in December 2025 and by March had 84 percent of engineers using it (with 70 percent of all committed code originating from AI, the highest publicly reported ratio at any major technology company) it announced it had AI opening 11% of new pull requests, which reveals precisely what percentage of committed code came from AI and nothing whatsoever about whether that code was worth what it cost to produce.
CTOs built internal dashboards ranking engineers by AI consumption, which functionally made them cost-acceleration mechanisms. Like Amazon, who in the same period it was announcing 16,000 layoffs, launched a token consumption rankings page, Meta called theirs “Claudeonomics” and handed out badges like “Token Legend” and “Session Immortal” — the cobra farm stated as corporate policy, formalized into vocabulary, and gamified into a leaderboard. Uber also built internal dashboards ranking engineers by AI consumption and found their top engineers were burning between $500 and $2,000 each per month. Uber’s full-year AI budget was exhausted by April. Microsoft, after a quarter of telling investors that Lloyds Bank was saving each employee 46 minutes per day, quietly canceled most of its internal Claude Code licenses in May 2026, effective June 30, directing its own engineers to its own cheaper tool. And Bryan Catanzaro, Nvidia’s own VP of Applied Deep Learning, told Axios that for his team, the cost of compute had already exceeded the cost of employees. Taken together, these findings that explain why the earnings call slides require such careful construction.
The retcon goes: layoffs are transformation, headcount reductions are investments, undemonstrated gains are early signals, companies without ROI are simply early on the curve, and the curve bends upward: number go up.
Because large language models remain structurally prone to hallucination and architectural regression, cutting the human layer doesn’t eliminate the labor cost, it forces the enterprise to pay twice and both bills are on the OpEx side. Once for the token bills of the autonomous agents, and again for the specialized engineers whose job is to untangle what the agents produced. Salesforce will spend nearly $300 million on Anthropic tokens this fiscal year, against a global engineering payroll of roughly $5 billion for 15,000 engineers. That the token bill is no longer a rounding error in an R&D budget but a massive, recurring, variable operating cost that scales with every line of automated output does not make the financial decision any less rational or the amount any less commensurate, and it also does not replace the humans required to verify that output. It joins them on the invoice, at a rate that makes the original headcount reduction look, in retrospect, like a discount. Which it kind of was.
The Gartner data on who succeeds points elsewhere. The companies doing best are not the ones that replaced humans with AI or maxed out their token spends. They invested in people alongside it, building systems where humans supervise and extend what AI produces (which is not to imply Benioff isn’t doing that, I don’t have that insider info.) It’s a narrative that’s easy to nod along to and not so easy to put into practice. And even when it is put into practice, the bill has a way of coming due anyway, as we’re seeing with current obsession over the token cost reduction.
The “TIP” (Token Improvement Plan)
It starts as a dark joke inside engineering Slack channels, usually right after an unhandled recursive test loop burns through five figures in API calls over a long weekend: “Management is putting you on a TIP—a Token Improvement Plan.”
But as the inverse economics of the frontier solidify, the reporting layer is already building the infrastructure to make it real.
The conversation happens behind a closed glass door, guided by a Director of Engineering staring at a Datadog billing visualization:
“Your velocity is fine, but your token consumption exceeds twice your base salary. You’re appending a 100k-token repository to every prompt just to fix minor layout bugs. We need you to drop your marginal token spend by 45% over the next thirty days, or we’re going to have to route your IDE access through a smaller, open-source model running locally on an older Mac Studio.”
This is the ultimate loop of the “Doing Less with More” doctrine. The enterprise pink slips 15% of the human staff to fund the autonomous transition, only to discover that the remaining humans consume tokens like a runaway process, like a horse fitted with a jetpack (which is obviously cool.)
Doing more with less was the promise. Doing less with more is the condition: less verifiable return, more spend; less human understanding of the systems humans are nominally managing, more dashboards measuring the proxies that replaced that understanding; less signal, more noise; less stability, more recovery. Goldman Sachs forecasts a 24-fold increase in enterprise token consumption by 2030. (Let that sink in.) Gartner projects that even as individual token prices fall 90 percent, total enterprise AI costs will increase, because agents consume tokens at rates that make price-per-token beside the point. (Opus 4.7 consumes 4× as many tokens for the exact same prompts.) The math does not improve at scale; it compounds. The invoice arrives later and larger, and the people who receive it are not, in many cases, the people who signed the contract.
The story currently circulating is as clean as retcons get: AI is working, the adoption curve is healthy, the returns are coming, the companies that cut too deep are simply early, and the capital is flowing in the right direction. What capital is actually building is a different question. And unlike the productivity claims, unlike the 46 minutes saved at Lloyds and the 45 percent gains at IBM and the budgets blown away by April, and the dashboards and the tokenmaxing and the shadow IT cobras running under everyone’s desk; the answer is not a matter of measurement lag or methodology.


