Skip to content
← All articles
AI TransformationAugust 2, 2026 · 10 min read

Execution Is Getting Cheap Faster Than Verification Is

Individual productivity jumps after an AI rollout. Organizational velocity doesn't. The gap between those two curves has a name, and most firms are quietly cutting the roles that close it.

By Bharat Sharma
Share

A pattern is showing up across enterprises: individual productivity jumps after an AI rollout. Code ships faster, analysis that took a week takes an afternoon. Organizational velocity barely moves. Decisions still queue. Strategy still takes as long to execute.

You can watch it happen at a single desk. Someone who approved five pieces of work a day now receives fifty, each one polished and confident, with the same authority to sign, the same liability for signing, and the same number of hours in the day.

The reflex is to call this an adoption problem: better prompting, better tooling, more training. It's something more specific. The cost of producing work has fallen sharply, and the cost of establishing that the work is correct and safe to act on has not fallen nearly as fast. Where verification governs release, where nothing ships, posts, or gets filed until someone signs, the organization runs at the slower curve.

Call the accumulating difference verification debt: the gap between what an organization can now produce and what it can still stand behind. That gap is the thing worth planning around. Not whether AI flattens the org chart, a claim that became conventional wisdom in about eighteen months on evidence that wouldn't survive serious review.

Verification isn't one thing, and that's the crux

The obvious objection to any "verification is the bottleneck" argument: why wouldn't AI verification get cheap too? Automated testing, evaluation harnesses, red-teaming, monitoring, agents auditing agents. All improving fast. The answer requires taking verification apart. The useful cut isn't by activity but by what makes each part hard.

Checking is hard because criteria are tedious to apply at volume. The criteria themselves are already written down. Tests pass, figures reconcile, the citation exists, the clause is present. Anything whose difficulty is volume-against-stated-criteria automates well, and this is automating quickly.

Judging is hard because the criteria aren't written down. Is this the right answer to the actual problem, given context nobody documented? Is the model answering a subtly different question than the one asked? Is this technically correct and strategically stupid? Difficulty here comes from unstated context, which is why it automates unevenly and degrades precisely on the novel cases that matter most.

Underwriting is hard because someone bears the consequence. Not who checked, but who answers for it in front of a regulator, a court, a customer, or a board. Difficulty comes from consequence-bearing, which is a property of legal and social standing, not of capability.

Three categories because there are three distinct sources of difficulty: volume, context, and consequence. Each responds differently to automation, which is the entire point. You could subdivide further, but any finer cut splits things that behave the same way under the same forces.

Aviation shows the split cleanly: nearly every inspection is automated, and a human still signs off. That isn't ceremony. Checking scaled; authorization didn't.

So: checking is automating fast, judging is automating unevenly, underwriting isn't automating at all. As output volume rises, the second and third become the constraint. Firms that instrument only the first will conclude their capacity is fine right up until it isn't.

One honest limitation before going further. This is a mechanism argument supported by industry examples and institutional analogy. Nobody has measured the relative slopes of these two curves across sectors, and I certainly haven't. If you want the empirical version, it doesn't exist yet.

What this looks like when it breaks

Picture a commercial credit officer who reviewed five credit memos a day, each drafted over hours by an analyst whose reasoning she could interrogate. Her team now generates fifty, each polished, internally consistent, and confident. Her sign-off authority hasn't changed. Her liability hasn't changed. Her day hasn't changed.

She will not review fifty memos. She will develop a heuristic for which ones to actually read, and approve the rest on the strength of how they look.

Human factors research has a name for the mechanism: automation bias, the tendency to accept automated output more readily as it becomes more fluent and more voluminous, particularly under time pressure. Polished, confident, high-volume output is precisely the condition that produces it. AI has made fluency free.

That's the failure mode. Not robots making mistakes. Rubber-stamping under liability: humans nominally accountable for volumes they can't inspect, in a system that still records their signature as if they had.

Nothing about this is software-specific. The shape recurs wherever output volume is rising, errors are consequential, and expertise sits in a shrinking senior layer: radiology reads, claims adjudication, audit sampling and workpaper review, regulatory document preparation in pharma, credit memos, asset inspection in energy, contract review, clinical documentation.

And one thing to be clear about, because it's the obvious misreading: the answer is not to push senior people to review faster. That accelerates the failure. Throughput at the sign-off layer is the thing you are trying to protect, not optimize.

The second-order effect

The work being automated first is entry-level and mid-tier cognitive execution: precisely the work that historically built the judgment needed to verify. Reviewing the routine cases is how people learned to recognize the non-routine ones. Cut those roles and you erode the pipeline that produces senior verifiers. Call it the Missing Junior Loop.

The support here is structural rather than statistical. Every profession that carries serious consequences for error (medicine, aviation, law, accounting, the skilled trades) has independently converged on mandatory supervised hours before independent authority, and has kept that requirement through decades of automation that made the underlying work far easier. Those institutions are not sentimental about training costs. They maintain apprenticeship because the alternative has been tested and rejected where failure is legible.

That's an argument from institutional convergence, not evidence that removing apprenticeship in knowledge work degrades expertise on a specific timeline. Nobody has that evidence; the experiment is running now.

The serious counterargument is that simulation and synthetic practice can build judgment faster than legacy junior work did, deliberately concentrating hard cases instead of waiting for them to arrive. That may well be true for pattern recognition. Where it likely falls short is that judgment appears to require consequence: the memory of having been wrong when it was yours, and having to say so. Simulation reproduces the case, not the exposure. A junior who has never owned an error has learned what mistakes look like without learning what it costs to make one.

I won't manufacture urgency about timing. Whether this bites in five years or fifteen, the asymmetry favors acting. Maintaining apprenticeship capacity you didn't strictly need is cheap; discovering in a decade that you can't produce senior judgment is not.

The org-chart story, held loosely

There's a tidy structural story that tends to ride along with this one. As protocol coordination replaces human coordination, the firm drifts toward an hourglass: a thin senior layer holding intent and accountability, a protocol layer encoding routing and policy, abundant agent execution at the base.

It's a plausible model resting on thin evidence. The measured effects are confounded by interest rates and post-pandemic overcorrection, and the counterexamples are serious: Amazon is AI-intensive and deeply hierarchical, Apple runs a functional hierarchy with extreme compartmentalization. In high-reliability domains like aviation, nuclear, and clinical care, AI adoption plausibly deepens hierarchy, because auditability and clear lines of authority are the design goal.

Test the hourglass in specific value streams; don't reorganize around it. The verification argument deliberately doesn't depend on it, so the durable part isn't hostage to the fashionable one.

The thing that actually kills these initiatives

Organizations don't only coordinate work. They allocate status, budget, promotion, and power. AI barely touches any of it. So the governance layer everyone recommends building is not a technical project. Call it an AI Orchestration Office or don't; the function is ownership of workflow architecture, agent operations, and policy encoded as executable guardrails rather than PDF. Either way, it is a claim on authority.

Whoever writes the routing rules decides which work is legible, which team's process becomes the default, and whose exception gets auto-approved. That's the authority existing function heads hold today, and they won't concede it because an architecture diagram says so. Protocols are never politically neutral. Every transformation that treats a control plane as a tooling decision discovers this around month nine.

Worth noting: flattening tends to intensify political competition rather than resolve it. Fewer rungs means the same ambition contests fewer positions.

What I'd actually do

Instrument verification capacity before you cut headcount. Most functions know their throughput and not their audit bandwidth. Four questions: who signs off, how long does a competent review take, what happens to that number at 10x volume, and what fraction of current sign-offs are genuine reviews versus rubber stamps. The last one is uncomfortable and the most informative. This is a diagnostic, not a formula. I'd rather give you honest questions than a metric with invented precision.

Design the three activities separately. Automate checking aggressively. Instrument judging: track where human reviewers overturn agent output, because that's your live map of where judgment is still load-bearing, and where the overturn rate approaches zero you have either a solved problem or a rubber stamp. Never let underwriting stay implicit; if nobody can name who bears the consequence, you have an unowned risk, not an efficient process.

Treat apprenticeship as infrastructure. Structured human-in-the-loop review where juniors verify agent output under supervision, deliberate exposure to failure with real if bounded consequence, rotation through the systems they'll eventually sign for. I'd do this even if everything else here is wrong.

Pay for underwriting, not output. If the scarce capability is bearing consequence, and promotion and compensation still reward volume produced, you are selecting against the exact thing your sign-off layer is about to need. And you'll lose the people carrying the most risk first, because they're the ones paid least well relative to what they absorb.

A note for readers without the authority to change headcount or compensation. You can still run those four questions against your own team, and the answer is the most useful thing you can put in front of someone who does have that authority. "My team can now produce fifty of these a week and I can competently review twelve" is a specific, checkable claim about capacity, not a complaint about workload. The instrument works best held by the people living with the constraint. Protecting your own juniors' exposure to real cases, with real if bounded consequence, is the other thing that needs nobody's permission.

What would prove this wrong

The argument breaks if judging becomes reliably automated on novel cases; if synthetic practice demonstrably produces judgment equivalent to consequential experience; if liability genuinely decentralizes; or if AI liability insurance and vendor indemnification mature to the point that firms transfer real exposure to third-party underwriters rather than named individuals. That last one is the most plausible of the four and the one I'd watch hardest. Insurance markets are good at pricing risks nobody wants; if they price this one, the underwriting bottleneck loosens considerably.

Absent those, the asymmetry is the thing to plan around.

And the argument isn't really about AI. It's about what happens to an organization when execution scales faster than judgment and accountability, when verification debt goes unpaid long enough to come due. AI is the first technology to open that gap this wide, not the last. Which means the answer outlives whatever the current generation of models turns out to be capable of.

Execution is getting cheap faster than verification is. Most organizations are currently cutting the exact roles that produce their future verifiers. You can disagree about the shape of the org chart and still have that problem.

#AI#Engineering Leadership#Operating Model#Governance