Verification Tax
Verification Tax
Steven Brovich's named force in his four-tensions framework (see A Leaders Guide to Advanced Team Structures (AWS Events)):
"AI generates code 10× faster, but it's 3× harder to validate. The reviewing bottleneck eats the velocity if you're not careful."
warning Stat traced (web verification, 2026-07-07): no published study matches the 10×/3× multipliers — treat Brovich's numbers as illustrative, not empirical. The closest published evidence is Faros AI's AI Productivity Paradox Report (Jul 2025; telemetry from 10,000+ developers, later 22,000): high-AI-adoption teams merge 98% more PRs but PR review time rises 91% (PR size +154%, bugs/dev +9%), with no improvement in company-level delivery metrics — i.e. roughly "2× output, 2× slower review." The tax is real; the multipliers are smaller and better-sourced at ~2×/2×.
Why it's a distinct concept
The vault already tracks Hallucination Laundering (Keen's accountability frame — submitting plausibly-confident AI output as your own) and Binary Eval Assertions (Karpathy's auto-research-loop mechanism for what makes verification cheap). The verification tax is the operations version: even when humans are diligently verifying, the verify step takes 3× longer than the generate step. So 10× generation speedup nets out far less than 10×.
The implication is the design problem the loops as core primitive page implicitly grapples with: the verification step must itself be made cheap (binary evals, automated tests, codable assertions) or the apparent speedup is illusory.
Connects to two big vault themes
- The Binary Eval Assertions argument is the engineering answer to the verification tax. When verification is a binary true/false assertion that can be machine-checked, the tax falls to ~0. When it requires human reading, the tax is real and Brovich's 3× number is plausible.
- The AI Productivity Disconnect is the macro-empirical version of the same gap. The Bank of Korea 2026-12 paper found AI saves 3.8% of work time but worker-level output correlation is zero. One mechanism for the disconnect: the time saved on generation is given back to verification — and on average, in current workflow design, more than given back.
Held in tension with three other forces
Brovich names verification tax as one of four simultaneous forces the leader must hold in tension:
- Expert multiplier — senior people with AI are order-of-magnitude faster (Project Mantle)
- Bottleneck shifts — not "can we build it?" → "do we have the data and can we decide fast enough?"
- Verification tax (this page) — 10× faster generation, 3× harder validation
- Deskilling trap — juniors faster + less grounded
"All four are true at once. The leader's job is to hold the tension."
Cross-references
- Botsitting — the survey-measured knowledge-worker-wide version (Glean: 6.4 hrs/week making AI usable); broader than verification — includes context-feeding pre-work and rerun re-work
- A Leaders Guide to Advanced Team Structures (AWS Events) — canonical (Brovich)
- Binary Eval Assertions — the engineering antidote (make verification cheap)
- Hallucination Laundering — what happens when the tax goes unpaid (Keen)
- Auto Research Loop (Karpathy) — the loop primitive that lives or dies on cheap verification
- AI Productivity Disconnect — the macro-empirical evidence the tax is real and not yet engineered around
- Effective Feedback Compute — the scaling-laws framing of how much informative verification you get per unit compute
2026-07-21 — Uber's operating-at-scale evidence
How to Help People Thrive with AI (AI Daily Brief) surfaces the first vault-visible enterprise data on operating at scale with the verification tax paid. Praveen Napali (Uber CTO) reports:
- 70%+ of Uber PRs attributed to local or cloud agents
- 2,500+ agent skills built across the software development life cycle
- 99% of engineers use AI tools
Uber has not disclosed its verification approach (how PRs are reviewed, what evals exist, how skills are QA'd), but the combination — high agent-attributed PR rate + shipping-rhythm sustained + business-function pods graduating in 2-week timeboxes — is empirical evidence that the tax is payable at scale in a large enterprise. This is the first vault-visible existence proof; the mechanism is still unknown and worth watching for.
The Agentic Pods programme adds an important second point: pod validation is built into the 10-day cadence — days 6-9 ("Validate with several others performing the same work; does it generalise; does it actually make their job better?") is a formalised verification phase inside the pod's own workflow. So Uber has at least built verification into the skill-authoring loop, not just the code-review loop.
2026-08-09 — agent-checks-agent as the emerging engineering answer
Nate B. Jones in Your Chatbot Hallucinated in 2024 Your Agent Lies in 2026 (Nate B Jones) surfaces a specific engineering pattern for driving the tax down on the agent-lying class specifically: a separate agent whose only job is to inspect the working agent's tool-call requests for intent-alignment, mid-execution. Both Claude Code and Codex have this shipped (Jones's transcript wording: "approve forming or review forming").
This is a specific implementation of the Binary Eval Assertions answer applied at the action layer rather than the output layer:
- Output-layer verification (existing pattern): check the diff / the artefact after the fact — expensive because you're reasoning about the finished state
- Action-layer verification (new pattern): check each tool call as it's requested — cheaper per-check, catches the substance-vs-shape divergence at the point it happens, and the reviewer agent has a much smaller decision surface than a full post-hoc code reviewer
The tax remains real. But the shape of the tax changes: instead of paying it in human review-time on completed PRs (Faros AI's 91% review-time inflation), you pay it in automated action-time supervision that runs continuously and cheaply. This is the same primitive Ryan Lopopolo's persona-shaped reviewers implement, moved from the CI-on-push insertion point to the tool-call insertion point.
The design implication for anyone building at Uber-scale (70% agent-attributed PRs, 2,500 skills): the answer to the tax at scale is not more human reviewers — it is layered automated review, with human attention concentrated where the automated layers disagree. Verification-Tax mitigation stops being a code-review-workflow problem and becomes an agent-architecture problem.
2026-09-26 — the human residue: cognitive debt and AI fatigue
Full Course Spec-Driven Development with Coding Agents (DeepLearningAI) names the felt side of the tax for individual developers: AI fatigue ("so much to review") and Cognitive Debt — the growing gap between what the agent has written and what the human understands. Its practitioner mitigations are workflow-level rather than architecture-level: small diffs with frequent commits, clean breaks between features, review at the level of spec intent (not variable names), reading tests under a debugger as comprehension, and a subagent deep-review pass before merge — an output-layer instance of the agent-checks-agent pattern above. Useful framing distinction: the verification tax is a throughput cost (review time); cognitive debt is the comprehension residue when that cost is skimped.
2026-09-26 — skill-embedded verification (Anthropic-engineer pattern)
What to Build Instead of AI Agents (Nate Herk) surfaces the individual-skill answer to the tax, attributed to Anthropic engineers Barry Zhang + Mahesh Murag:
"A skill shouldn't hand you its first attempt and call the job done… Your first look should not be the agent's first look. It should be the agent's fourth or fifth or sixth look."
The concrete pattern: encode the check inside the skill — for a slide deck, render each slide, screenshot, inspect, fix cropping/readability, rerender; for a research report, open the primary sources, match claims to evidence, remove anything unverifiable; for persuasive copy, run the draft past persona subagents (beginner / skeptical-buyer / target-audience). Non-negotiable rule: verification needs evidence outside the first draft — a screenshot, a test, a source, a reference example, or persona feedback — not the same model reading its own output and approving it. A reusable skill-embedded prompt is on the source page.
This is a third insertion point for verification-tax mitigation, complementing the two already tracked on this page:
| Insertion point | Pattern | Vault anchor |
|---|---|---|
| Output layer (post-hoc review) | Human PR review, subagent deep-review pass | Full Course Spec-Driven Development with Coding Agents (DeepLearningAI) |
| Action layer (per tool-call) | Reviewer agent checks tool calls against intent mid-execution | Your Chatbot Hallucinated in 2024 Your Agent Lies in 2026 (Nate B Jones) |
| Skill layer (self-check inside the skill) | Skill runs, inspects its own output against pre-defined acceptance criteria, iterates until met | What to Build Instead of AI Agents (Nate Herk) |
The skill-layer answer is cheaper per invocation than the action-layer answer (no separate supervising agent per tool call) but narrower — it only catches problems the skill author already anticipated. Best paired with the other two layers, not used alone. Directly applicable to this vault's /ingest, /query, /lint, /journal, /crm operations: their implicit binary invariants (index updated, log entry well-formed, raw file moved) should become explicit pre-return checks inside the skill.
2026-09-27 — pay the tax selectively: blast-radius review
How I Review AI Code (John Kim) adds an allocation answer to the three insertion points above: don't pay the tax uniformly. John Kim (senior staff engineer, Meta) scales human reading to the change's blast radius — deep review for trunk code (shared state, core infra, entry points), a skim plus agent-produced proof for leaf code (isolated or feature-gated). Feature gating is what makes most changes leaf-like, because it makes them reversible. See Blast-Radius Code Review. His other levers map onto this page: proof-carrying PRs (tests, logs, screenshots, agent confidence levels) are output-layer evidence; a fresh-context adversarial reviewer agent is the output-layer agent-checks-agent pattern; nits go to linters. Caveat he's explicit about: the approach gets you merge-ready, not launch-ready — the last ~20% (whole-feature audit, refactor, polish) still costs roughly as much human time as the first 80%.
Sources
- A Leaders Guide to Advanced Team Structures (AWS Events) — canonical
- Full Course Spec-Driven Development with Coding Agents (DeepLearningAI) — 2026-09-26; cognitive debt / AI fatigue and developer-level mitigations
- How to Help People Thrive with AI (AI Daily Brief) — 2026-07-21; Uber operating-at-scale evidence
- Your Chatbot Hallucinated in 2024 Your Agent Lies in 2026 (Nate B Jones) — 2026-08-09; action-time supervision as engineering answer
- What to Build Instead of AI Agents (Nate Herk) — 2026-09-26; skill-embedded verification as third insertion point
- How I Review AI Code (John Kim) — 2026-09-27; blast-radius review as allocation answer