What to Build Instead of AI Agents (Nate Herk)
What to Build Instead of AI Agents (Nate Herk)
Nate Herk's read on a public account by two Anthropic engineers — Barry Zhang and Mahesh Murag (transcript renders as "Barry Jien" and "Mahesh Marog"; corrected here — Zhang leads AF at Anthropic and Murag is the MCP/skills-side engineer widely cited alongside him) — that Anthropic "basically stopped rebuilding a separate agent for every single job because the agent underneath had become way more general purpose than they expected." The unit of AI capability at Anthropic has moved from agent to skill loaded onto one general-purpose agent.
Video framing analogy (worth keeping verbatim as a leadership frame):
"The model is like the processor. The agent runtime is like the operating system. And skills are the apps."
So Claude Code already reads files, writes code, calls tools, and works through a task — you don't rebuild that stack for every new job. You author a skill that carries the process, context, scripts, and examples for the specific job and hand it to the same general-purpose runtime.
The four practical techniques
Nate's condensation of the Anthropic engineers' account into four operational rules for making skills work in practice.
1. Save proven code — don't rediscover the solution every time
Their team kept watching Claude write basically the same Python slide-styling script every run, burning tokens and producing non-deterministic results. Fix: save the script inside the skill "as a tool for its future self." The next run executes the proven file rather than regenerating the code.
Developer-idiom name: DRY (Don't Repeat Yourself). Applied to agents: if you already have a solution in code, don't pay the model to rediscover it.
Working prompt they recommend:
"Save the script you just used inside the skills script folder. Update the skill.md so future runs execute that file instead of trying to rewrite it. Then run the skill again and verify the result."
Testing rule they attach: run the same task twice, compare the important parts, confirm the skill is actually calling the saved file — the surrounding AI output may still vary, but the load-bearing piece is now deterministic.
Vault connection: this is the same primitive that makes loops cheap — replace "fresh guess" with "proven code" at every step you can. Directly relevant to the vault's own skills — the ingest/query/lint/journal/crm operations codify workflow; the pattern says also codify the scripts they call.
2. Progressive disclosure + precise descriptions
Analogy: a mechanic owns hundreds of tools but doesn't dump all of them on the bench to change a tire. Skills work the same way — Anthropic calls it progressive disclosure.
- Claude starts with only the YAML frontmatter (name + description) of each skill, not the full instructions or scripts.
- When a user prompt matches a skill's description, that skill's
SKILL.mdgets read in full. - Larger reference files and scripts stay in the folder until the task actually needs them.
This is what keeps irrelevant instructions out of working context and avoids the bloat sometimes called context rot. But it depends entirely on description quality:
- Weak: "help with content" vs "create marketing assets" — overlapping and vague.
- Strong: "This skill creates LinkedIn carousels from a topic, from a transcript, or an outline. Use this when the user asks for a carousel, carousel slides, or a LinkedIn document post." — Claude now knows what the skill does and when to use it.
Rules he lifts from the Anthropic account:
- One specific job per skill.
- Description uses the words a real user would say.
- No two skills competing for the same request.
Audit prompt he recommends (usable on the vault's own skills):
"Review all my skill descriptions. For each one, tell me what it does, when it should trigger, and where it overlaps with any other skills. Rewrite only descriptions that are ambiguous."
Three-test verification per skill: (a) obvious request that should trigger, (b) differently-worded request that still should trigger, (c) unrelated request that must not trigger.
"A skill that Claude can't find is basically a skill that you don't have."
Vault connection: reinforces the Skills page's existing description-quality thread and Pocock's Trigger dimension in Skill Checklist (Pocock). The three-test cadence is a small, adoptable pattern this vault doesn't yet formalise on its own skills.
3. Turn corrections into durable instructions
The lesson every user throws away: when you correct Claude and close the chat, that correction is gone.
Anthropic's stance per Nate: "anything Claude writes down can be used efficiently by a future version of itself." Skills are the durable place. Rule of thumb:
- Process was wrong → update
SKILL.mdinstructions. - Missing voice / brand / examples → add a reference file.
- Same mistake keeps happening → add a clear rule that prevents it explicitly.
Then rerun the same task to verify. The prompt shape he suggests:
"Review what went wrong during this run. Decide whether the cause was the process, missing context, a weak rule, or unreliable code. Update the skill in the smallest durable place. Then rerun the same task and verify the fix."
His own AIOS example: when an agent said it couldn't find a file he knew existed, he doesn't hand it the path and move on. He asks it to backtrack, show its work, figure out why it missed the file, and update the routing or the skill so the next run starts in the correct place.
Honest hedge he adds: "model-proof" is not literal. Different models still interpret skills differently, but the process becomes portable — agent skills are an open format and the same skills folder can work across compatible harnesses (he names Codex and Hermes explicitly). Test important skills against another compatible agent; if it falls apart, look for hidden assumptions or model-specific instructions.
Vault connection: this is the manual counterpart of the overnight auto-research loop on the vault's Skills (Claude Code) page — small, human-in-the-loop refinements after every use, made durable in the smallest place. The rehearsal ritual ("update the smallest durable place, rerun, verify") is the practical unit.
4. Bake the verification into the skill
"This one is probably the most important. A skill shouldn't hand you its first attempt and call the job done."
The gap most workflows have: skill runs → file saved → agent says "done" → you open it → formatting is broken, sources don't back the claims, script falls flat with the target audience. The AI did 70–80% of the job; you do the last 20–30%. If you already know how to check the work, encode that check inside the skill.
Concrete verification patterns by output type:
- Slide deck → render each slide as an image, inspect screenshots, fix anything cropped/hard-to-read/out-of-bounds, rerender.
- Research report → open the primary sources, match claims to evidence, remove anything unverifiable.
- Script / ad / persuasive copy → run it past persona subagents (beginner agent flags confusion; skeptical-buyer agent flags disbelief; target-audience agent flags click-away moments). Don't accept every note — but issues raised by more than one persona are strong signals.
Non-negotiable rule stated explicitly:
"Verification isn't Claude reading its own work and saying 'looks good to me.' It needs some kind of evidence outside of that first draft — a screenshot, a test result, a source, a reference example, or feedback from different agent perspectives."
Reusable skill-embedded verification prompt:
"Before returning the final output, define the acceptance criteria. Create the first version. Inspect it using the relevant verification method. Fix every issue you find and then run another pass. Return the output only after it meets the criteria with a short summary of what you checked. If something can't be verified, tell me exactly what remains."
Better still: an objective success metric the agent keeps working against until it hits.
"Your first look should not be the agent's first look. It should be the agent's fourth or fifth or sixth look."
Vault connection: this is the Verification Tax mitigation pattern moved inside the skill rather than left as a downstream review step — see the 2026-09-26 update on that page. Structurally similar to Nate B. Jones's action-layer supervision (from Your Chatbot Hallucinated in 2024 Your Agent Lies in 2026 (Nate B Jones)) but at the output layer, and self-contained inside a single skill. Also aligns with the auto-research pattern's need for cheap verification: the persona-subagents pattern is a soft-eval addition on top of binary assertions for cases where the check is subjective.
Why this matters for the vault
Three load-bearing takeaways for someone running an AIOS on top of the Second Brain pattern:
- The unit of AI capability is the skill, not the agent. Reinforces the Skills (Claude Code) thesis and Skill Engineering as a discipline; useful executive-audience phrasing is the model/OS/apps analogy above.
- Every skill has four failure modes and four operational rules that map onto them: rewriting solved problems (fix with saved scripts), unclear activation (fix with precise descriptions + progressive disclosure), lost corrections (fix with durable instruction updates), premature "done" (fix with skill-embedded verification). This is a skill lint rubric that maps cleanly onto Pocock's Trigger/Structure/Steering/Pruning but adds an explicit Verification dimension Pocock's rubric under-treats.
- Verification is a skill concern, not a workflow concern. This is the specific engineering answer to the Verification Tax at the individual-skill granularity, complementing Uber's pod-level answer and Nate B. Jones's action-layer answer.
Key claims to challenge later
- The "Anthropic engineers stopped building agents" framing is a YouTube-headline compression of "stopped rebuilding a separate agent for every single job." The distinction matters: they didn't abandon the agent primitive, they abandoned the bespoke-agent-per-task pattern.
- Nate presents these as Anthropic-blessed patterns, but the source is his read of a public account by Zhang and Murag — the exact wording of the underlying talk/post is not quoted verbatim. If the original Zhang/Murag artefact surfaces (blog post, conference talk), pull it into this page as a primary source and downgrade this page's role to practitioner interpretation.
- The "model-proof" hedge is worth respecting: skills authored under one model regime silently over- or under-instrument the next. See Tyler Brown's middle-school-to-high-school framing on Skills (Claude Code).
Sources
- Video: https://youtu.be/HIRDzMtuWFk — Nate Herk | AI Automation, 2026-09-26 capture
- Trigger: source (Telegram Daily Learning #4240, transcript pre-enriched in capture)
Cross-links
- Skills (Claude Code) — the vault's canonical skills concept page
- Skill Engineering — the discipline
- Skill Checklist (Pocock) — the vault's judgement rubric; this video's four rules extend it with an explicit verification dimension
- Verification Tax — the tax this pattern pays inside the skill rather than downstream
- Auto Research Loop (Karpathy) · Binary Eval Assertions — the automated counterparts of technique 3
- Nate Herk (AI Automation) · Anthropic · Barry Zhang · Mahesh Murag