How I Review AI Code (John Kim)
How I Review AI Code (John Kim)
John Kim — a senior staff software engineer at Meta who runs an AI-tutorials YouTube channel as a side project — gives his working answer to the split in the engineering community over whether you should still read AI-written code. One camp says reviewing your own code means you're not moving fast enough; the other says agents still make dumb mistakes and ship AI slop. His answer: "it depends" — review is a gradient, and the depth should track the blast radius of the change.
Key claims
- The volume problem is real. One engineer running several agents generates more code than they can review — even their own, let alone teammates'. This is the Verification Tax seen from the individual contributor's side.
- Code base as a tree → Blast-Radius Code Review. Trunk code (entry points, shared state such as reducers, core infra like image rendering or networking) can take down the whole app, so he reads it closely. Leaf code (isolated components, new endpoints, anything fully behind a feature gate) gets a skim, and agent-produced proof (component, snapshot or unit tests) stands in for line-by-line reading.
- Plan for gating from the start. Gate every new feature (a Boolean flag is enough to start; LaunchDarkly-style feature toggling at scale). Separate leaf work from integration work and gate the integration layers. Reversibility is what buys permission to review less. His safety questions: what is the control experience that must not change? how bad is it if this goes wrong? can I roll it back? One-way doors get more review.
- Make agents bring proof. A PR should carry evidence: unit tests (skim test names and asserts), runtime logs, screenshots/video of the UI, and an agent-stated confidence level per part of the change. "Agentic validation is a core concept of being able to read less code." Side tip: write a skill that stops agents writing useless tests.
- Nits are dead. Style preferences don't belong in review any more — leave them to type checking, linting and agent reviewers. Refactors are cheap once validation exists.
- Review with a fresh, adversarial agent. Use Claude Code's
/code-reviewor Codex's reviewer, but in a separate agent that lacks the authoring context — the authoring agent "will cheat" and anchor on its prior conversation. See LLM as Judge. - PR template. Give agents a template so PRs look alike and carry the important bits. Pet peeve: a PR summary and test plan longer than the diff itself.
- Automate the review loop. Codex offers a GitHub reviewer that runs on PR open (reportedly a model tuned for review). To avoid review ping-pong, set a recurring goal that "babysits" the PR — hourly, address reviewer comments, rerun local validation, update notes, resubmit.
- Merge-ready ≠ launch-ready. Working this way gets you ~80% done ("broad strokes"). The last ~20% — auditing the whole feature with agents for performance, bugs, security and spec-fit, refactoring, polishing — takes about as long as the first 80% and is where human taste lives. See Agentic Engineering.
- Launch through experiments. Even without statistical power for real A/B tests, canary or gated rollouts surface crashes and alerts fast. The goal is an "agentic garden" of tests, validation and reviewer agents, with gating as the undo button.
He still reads more code than before, but selectively, and expects to read less as models improve. Knowing your code base well makes this faster: you can spot quickly which changes touch dangerous zones.
Assessment
A practitioner's operating model, not evidence: single author, no numbers, demonstrated on a toy app. Its value is in the vocabulary (trunk vs leaf, proof-carrying PRs, merge-ready vs launch-ready) and in how cleanly it composes with patterns the vault already tracks — Verification Tax mitigation, fresh-context agent reviewers, and skill-embedded verification from What to Build Instead of AI Agents (Nate Herk). The "reportedly custom-trained Codex review model" claim is his, unverified.
Connects to
- Blast-Radius Code Review — the concept page this source anchors
- Verification Tax — risk-tiered review as a way to pay the tax selectively
- LLM as Judge — fresh-context adversarial reviewer
- Agentic Engineering — the last-20% taste phase
- Claude Code, Codex — built-in review tooling