← narwal.one/Second Brain
SecondBrain
Ask the Brain
Index/Conceptupdated Sat Sep 26 2026 08:00:00 GMT+0800 (Philippine Standard Time)

Spec-Driven Development (SDD)

spec-driven-developmentsddmethodologyai-codingagentic-engineeringtddcontext-rotadlcbrownfield

note 2026-09-26 addition — the hands-on procedure (DeepLearningAI × JetBrains course) Full Course Spec-Driven Development with Coding Agents (DeepLearningAI) (Andrew Ng intro, Paul Everitt teaching) is the vault's first operating manual for SDD rather than a survey or critique. Its loop: Project Constitution (mission / tech-stack / roadmap, drafted in an agent interview) → per-feature branch + plan / requirements / validation specs → implement on fresh context → human validation → replanning on its own branch (including writing skills to automate the process). What it adds to this page:

  • Brownfield is in scope — the agent reverse-engineers the constitution from an existing codebase + backlog; after that the loop is identical. Rebuts "SDD is greenfield-only."
  • Disciplined, not naive — every diff is read, tests are stepped through under a debugger, a subagent deep review catches what the main agent missed, and fixes go to spec and code together. It sits firmly on the right-hand column of the split table below.
  • Answers two listed failure modes: constitution decay → mandatory replanning phase with constitution changes on a dedicated branch; the human-side cost → named as Cognitive Debt / AI fatigue, managed by small steps and clean breaks between features.
  • Framework landscape extended — adds OpenSpec (propose → explore → apply → archive) alongside Spec Kit, and argues for adopting one then customizing with skills; portability via MCP + AGENTS.md + Agent Skills + Agent Client Protocol (ACP).
  • Open problem named: linking spec versions to the code they produced is "an evolving topic in the community" — the traceability gap enterprises will care about.

warning Contradicts Spec-Driven AI — BMAD SpecKit GSD Superpowers (Vimal Dwarampudi) (minor attribution): that page attributes SpecKit to Microsoft; the course calls it GitHub's Spec Kit. Reconciled — Spec Kit is published by GitHub (github/spec-kit), a Microsoft subsidiary; "GitHub" is the more precise attribution.

Spec-Driven Development (SDD)

Umbrella concept. The 2026 methodology in which the specification — not code, not documentation, not chat transcripts — is the source of truth throughout the software lifecycle. The AI compiles spec into implementation; when reality diverges, you edit the spec and regenerate. Distinct from Karpathy's "prompts as the program" only in scope: a spec is a versioned, structured prompt that persists across sessions.

Two vault sources take the positive framing (this is the emerging discipline); one takes the negative framing (this is vibe coding by another name). Both can be right depending on which shape of SDD is in play — see the disciplined/naive split below.

The claim (Dwarampudi)

SDLC optimises for human coordination. ADLC (AI Development Life Cycle) optimises for autonomous execution with human oversight. Both share an unresolved tension: how do you ensure what's built matches what was intended? Requirements documents (SDLC) go stale; prompting harder (early AI coding) is unreliable. SDD is the answer: make specs executable.

The canonical SDD pipeline (per Dwarampudi, drawing on the SpecKit shape):

Constitution → Specify → Plan → Tasks → Implement

With two structural additions:

  1. A feedback loop — requirements change → update spec → regenerate.
  2. A constitution rail — governance you set once, respected everywhere downstream.

Developers' primary job shifts from writing code to authoring and maintaining precise intent.

The counter-claim (Pocock)

Matt Pocock argues in Software Fundamentals Matter More Than Ever (Matt Pocock, AI Engineer) that when this reduces to "write a spec → LLM compiles it → don't read the code → edit the spec when broken" it produces a compounding failure — each recompile makes the code worse via software entropy. Bad code stalls AI (AI in a good codebase is extraordinary; in a bad one it can't find its way and produces plausible-but-wrong output). Specs-to-code is divestment from design.

See Specs-to-Code for the full critique.

The disciplined / naive split (how to hold the tension)

The tension collapses once you split naive SDD from disciplined SDD:

Dimension Naive SDD (Pocock's target) Disciplined SDD (Dwarampudi's mature setup)
Spec role Sole source of truth; code is disposable Source of truth, but code is read and reviewed
Design concept Lives in the spec (Pocock: it doesn't) Lives in dialogue; spec is one persistence surface
Context per session One long-running session Fresh 200k instance per task (GSD)
Verification Human eyeballs the running output TDD gates; delete-code-before-test (Superpowers)
Failure mode Compounding entropy → plausible-wrong Constrained by the discipline layers
Vault verdict "Vibe coding by another name." The emerging engineering practice; still single-outlet framing

Both Pocock and Dwarampudi actually agree that naive SDD fails. They disagree on where the fix belongs:

These aren't mutually exclusive. Pocock's inner-session discipline is what a well-authored SKILL.md or Superpowers phase produces at the moment of implementation; Dwarampudi's outer-session methodology is what stops the inner discipline from being silently violated when a subagent skips a step.

The 2026 frameworks (Dwarampudi's four)

Each names a distinct problem:

  • BMAD (Breakthrough Method of Agile AI-Driven Development) — 21+ specialised agent roles with handoffs; the agile team as metaphor. Planning discipline; auditable artifacts. Doesn't solve context rot; QA runs after code.
  • SpecKit (Microsoft) — executable spec as engine; constitution-first pipeline. Cross-interaction consistency. Doesn't enforce TDD.
  • GSD (Get Shit Done) — Lex Christopherson. Fresh 200k-token instance per task; parallel waves for independent tasks. Solves context rot (Christopherson's coinage for the observable Claude-quality degradation as context fills). Solves runtime, not planning.
  • Superpowers — Jesse Vincent. Methodology enforcement; 7-phase process; TDD-strict (any code written before a failing test is literally deleted); spec-approval and code-quality review gates. 170k+ GitHub stars in three months as an adoption signal.

Dwarampudi's composition thesis: BMAD + SpecKit for planning/governance; GSD + Superpowers for runtime discipline. Four layers, one methodology.

Where SDD sits in the wider vault frames

  • Software 3.0 — Karpathy's paradigm where prompts are the program. A spec is a versioned, structured prompt. SDD is Software 3.0 with git.
  • Agentic Engineering — Karpathy's ceiling-raising discipline. SDD-with-discipline is one systematisation of it.
  • Vibe Coding — the floor-raising cousin. Naive SDD collapses back into it (Pocock's argument); disciplined SDD is the escape (Dwarampudi's argument).
  • Context Engineering — SDD's constitution + spec are context-engineering artefacts at the project scope.
  • Loops as Core Primitive — GSD's fresh-instance-per-task is a specific loop shape; Cherny's "5–10 Claudes in parallel" is the same instinct without a name.
  • Design Concept (Brooks) — Brooks's ephemeral design concept lives between collaborators. SDD attempts to persist it as a spec. The Pocock argument is that this fails because the design concept is inherently ephemeral; the disciplined-SDD counter is that a well-authored spec plus Grill Me plus review gates comes close enough for practical work.
  • Verification Tax — Superpowers' TDD-strict rule and spec-approval gates are direct interventions on the verification tax; they shift it from downstream (humans catch mistakes at merge) to upstream (tests catch them at commit).

Failure modes to watch

Even with the discipline layers, SDD has predictable failure modes worth naming:

  • Spec creep — the spec accretes edge cases faster than the engineer's ability to reason about them holistically. The spec becomes bigger than a well-modularised codebase, and readers can't hold it in their head.
  • Constitution decay — the SpecKit-style constitution goes stale as the product evolves; nothing forces its re-adoption on downstream work, so it becomes cargo-cult text.
  • TDD as theatre — Superpowers deletes code-before-test but doesn't judge test quality. Tests can pass while missing the point (Reward Hacking).
  • Context-rot false comfort — GSD's fresh-instance-per-task solves within-session degradation but introduces its own tax: cross-task context has to be re-derived, and the plumbing (what state to pass, what to isolate) becomes a new source of subtle bugs.
  • Architect-vs-engineer split — SDD reads as clean and disciplined to architecture leadership (Dwarampudi's audience) and reads as ceremony to shipping engineers (Pocock's audience). Both readings are informative signal about what your team already lacks.

2026-09-26 — Pocock's own pipeline is spec-centred (AI Skills with Matt Pocock (The Pragmatic Engineer))

Pocock, the vault's main SDD critic, says he has "mixed feelings" about the term ("a strange term, it encompasses too much") and repeats his specs-to-code failure story (change spec → code gets worse; "you're not supposed to look at the code… it was garbage"). Yet his working pipeline is: Grill Me → spec (the "destination document"; formerly "PRD") → tickets (one per session, sized to the smart zone) → AFK implement loop, with Wayfinder when planning itself spans sessions. That places him squarely in the disciplined column above: the spec is a persistence surface for decisions reached in dialogue, not source code.

He also meets the "isn't this waterfall?" objection directly: the answer is cheap aggressive prototyping (3–4 versions, pick one) as part of writing the spec — and, via Grady Booch, waterfall's real failure was multi-year planning horizons, not planning itself. His sizing rule (align-after for small changes; grill for one-session-but-hard-to-reverse; Wayfinder for multi-session) doubles as an SDD when-to-bother heuristic, matching big tech's pre-AI PRD thresholds.

Cross-links

Sources