LLM as Judge
LLM as Judge
Pattern: use an LLM to evaluate output (code, text, or another agent's behavior) against criteria you'd otherwise have to encode by hand. The judge can be a plain LLM call (read this, did it follow the rule?) or an agent with tools (read this, run it, did the test pass?).
When it earns its keep
- The criterion is hard to express in code. Regex can check a URL prefix; "is this code idiomatic for our codebase?" cannot. The judge fills the gap.
- You're testing context, not code. Per Context Development Lifecycle (eval phase) — a code linter checks the generated code; an LLM judge checks whether the prompt/context successfully steered the model toward the intended output.
- Persona-shaped review. Per Harness Engineering (Ryan Lopopolo, AI Engineer) — Ryan's team encodes "front-end architect," "scalability engineer," "reliability engineer" personas into reviewer agents that run on every push. Each judge sees the diff through one lens and surfaces P2+ blockers.
Two failure modes to design around
- Non-determinism. Same input, different verdicts. Patrick's mitigation: run the eval N times and adopt error budgets ("must pass 4/5"). Don't gate CI on a single run.
- Bullying. Ryan's warning: if every reviewer-agent comment is a hard block, the implementation agent gets drowned in conflicting feedback and stalls. The right default is implementation agent can acknowledge, defer, or reject — judges propose, the implementer disposes.
Where this pattern shows up in the wiki
- Patrick Debois (Context Development Lifecycle) — judges as the testing layer for context itself
- Ryan Lopopolo (Harness Engineering (Ryan Lopopolo, AI Engineer)) — judges as reviewer agents in CI on every push, replacing synchronous human code review
- Karpathy (Andrej Karpathy on Agentic Engineering (Sequoia AI Ascent)) — adversarial-agent evaluation as the proposed hiring exercise: 10 codecs from another candidate try to break your Twitter clone
- Jones (Your Chatbot Hallucinated in 2024 Your Agent Lies in 2026 (Nate B Jones)) — the action-time variant: a separate agent that inspects tool-call requests and actions by the working agent for intent alignment, mid-execution rather than post-execution
2026-08-09 — action-time variant ("approve forming / review forming")
Jones names an emerging shipped-in-product variant of the judge pattern: rather than reviewing finished work after the fact, a separate agent inspects each tool-call request by the working agent as it happens, and can challenge or block it before the action lands. Both Claude Code and Codex have this implemented (Jones's transcript wording is "approve forming or review forming" — likely a mishearing of the actual feature name; the mechanism is real and shipped).
The action-time variant matters for a specific class of failure the batch-review variant misses: agent-lying failures where the artefact looks correct (an attached file with the right name, a passing test) but the state substitution happens mid-execution and would only be caught by inspecting the tool-call chain, not the final output.
Composes cleanly with Ryan Lopopolo's persona-shaped reviewers (batch, post-execution, on-push) — same primitive at a different insertion point in the loop. A team taking this seriously might run both: persona reviewers on push, action-time approver on tool calls, with an escalation path when the two disagree.
2026-09-26 — Non-LLM judges for online evals (Judgment Models (System 1 AI))
How People Are Actually Using Jev (AI Daily Brief): LangChain's Harrison Chase calls Jev "great for evals, especially online evals where you want to grade lots of traces." An Every test put numbers on the trade: Jev caught 6/7 planted writing errors in 0.35 s vs Claude Fable 5.1's 7/7 in 8.83 s, at ~580× lower cost. The pattern that emerges is tiered judging — a judgment model grades every trace against decomposed yes/no rubric questions (see Binary Eval Assertions); an LLM judge handles the flagged minority and anything needing reasoning. Caveat: judgment models return a score with no rationale, so they can't replace the LLM judge where the explanation is the deliverable.
2026-09-27 — The reviewer must not share the author's context
How I Review AI Code (John Kim): when using an agent to review agent-written code (Claude Code /code-review, Codex reviewer), run it as a separate agent without the authoring conversation — the authoring agent "will cheat" by anchoring on its own prior context. Practical corollary of the self-preference problem: the judge's independence comes from context separation, not just a different prompt. He pairs it with a PR template so the judge sees a consistent shape, and a recurring "babysit this PR" goal that fixes reviewer comments hourly until clean.
Sources
- Context Is the New Code (Patrick Debois, AI Engineer)
- Harness Engineering (Ryan Lopopolo, AI Engineer)
- Your Chatbot Hallucinated in 2024 Your Agent Lies in 2026 (Nate B Jones) — 2026-08; action-time variant
- How People Are Actually Using Jev (AI Daily Brief) — 2026-09; judgment-model tier below the LLM judge
- How I Review AI Code (John Kim) — 2026-09; fresh-context adversarial code reviewer