Skip to main content

Spec-Driven Development With Claude Code: A Full Walkthrough

Spec-driven development with Claude Code, step by step: a /spec interview skill, plan mode, a subagent plan review and hooks that gate each task on tests.

Founder of IImagined.ai

Published
Oct 11, 2026
Reading time
11 min read
Quick answer

Spec-driven development with Claude Code runs on three files and two gates. A /spec skill interviews you and writes SPEC.md, a fresh session in plan mode turns it into PLAN.md and tasks.md, and a read-only subagent reviews the plan. Hooks then freeze the spec and keep Claude from ending a turn while tests fail. GitHub's Spec Kit packages the same idea as ready-made skills.

Spec-driven development with Claude Code comes down to three files and two gates: a SPEC.md that Claude writes by interviewing you, a PLAN.md and tasks.md produced in plan mode and checked by a subagent, and hooks that stop Claude from editing the spec or ending a turn while tests fail. Every part is a documented Claude Code feature, so the walkthrough below needs no extra tooling. GitHub's Spec Kit, covered near the end, packages the same idea if you would rather install it than build it.

Checked October 2026 against Anthropic's Claude Code docs: best practices, skills, permission modes, subagents, the hooks guide, the hooks reference, commands and the tools reference, plus GitHub's Spec Kit quickstart and integrations reference. Claude Code ships often, so option labels and defaults can move between versions.

This is the hands-on companion to our Claude Code complete guide, which explains skills, subagents, hooks and plan mode one at a time. Here they are wired into a single workflow for a feature that is too big to describe in a sentence. For the method itself, including what belongs in each file and how much spec is enough, read our spec-driven development guide. Each step below ends with the result you should see before moving on.

The Claude Code spec workflow at a glance

The workflow moves a feature through files, and each file is written in its own session. Anthropic's best-practices page gives the reason: Claude's context window fills up fast and performance degrades as it fills. A spec on disk costs nothing to carry into the next session; a long conversation does.

From idea to reviewed diff
  1. 01
    Interview

    The /spec skill asks the hard questions and writes SPEC.md.

  2. 02
    Plan

    A fresh session in plan mode turns the spec into PLAN.md and tasks.md.

  3. 03
    Review the plan

    A read-only subagent checks both files against the spec.

  4. 04
    Gate

    Hooks freeze the spec and block the stop while tests fail.

  5. 05
    Execute

    One task per session, each with a check Claude runs.

  6. 06
    Review the diff

    A fresh subagent compares the result with the plan.

FileWhat it answersWritten byChanges when
SPEC.mdWhat the feature does, what is out of scope and how to verify it end to endThe /spec interview skillOnly when you change your mind, never mid-task
PLAN.mdWhich files change, and in what orderA fresh session in plan modeWhen the plan review finds a gap
tasks.mdNumbered units of work, each naming the check that proves itThe same planning sessionAs tasks are ticked off

The file names are conventions, not requirements. Anthropic's docs use SPEC.md and PLAN.md in their own examples, and Spec Kit uses lower-case spec.md, plan.md and tasks.md. Pick one set and keep it.

Before you start: the pre-flight check

The workflow only works if Claude has something to run. Anthropic's first best practice is to give Claude a check that produces a pass or a fail, because without one "looks done" is the only stop signal it has.

Pre-flight for a spec-driven feature
  • Claude Code is installed and logged in: claude --version prints a version
  • The repo has one command that fails loudly: a test suite, a build or a linter
  • CLAUDE.md names that command; run /init if the project has no CLAUDE.md yet
  • You are on a feature branch with a clean git status
  • jq is installed, because the hook scripts in Anthropic's docs parse JSON with it
  • The feature is too big for one sentence; smaller changes skip the spec

If the first item fails, our Claude Code install guide covers each operating system. Everything below works the same in the terminal and in the editor extension; the VS Code setup guide shows where plan mode and the diff review sit in that interface.

Step 1: Generate the spec with a slash command

Claude Code has no built-in /spec command, so you make one. Custom slash commands are skills now: a SKILL.md file in .claude/skills/spec/ becomes /spec. Save this as .claude/skills/spec/SKILL.md:

---
name: spec
description: Interview me about a feature, then write SPEC.md
argument-hint: [feature idea]
disable-model-invocation: true
---
I want to build $ARGUMENTS. Interview me in detail using the AskUserQuestion tool.

Ask about technical implementation, UI/UX, edge cases, concerns, and tradeoffs.
Don't ask obvious questions, dig into the hard parts I might not have considered.

Keep interviewing until we've covered everything, then write a complete spec to SPEC.md.

The body is the interview prompt from Anthropic's best-practices page, with $ARGUMENTS where the docs put a bracketed description. argument-hint shows the expected input during autocomplete. disable-model-invocation: true stops Claude from starting an interview on its own; the docs recommend it for workflows you want to trigger manually. Then run the command with your idea (the feature here is an illustrative example):

/spec a CSV export for the invoices page

Expected result: Claude asks multiple-choice questions through the AskUserQuestion tool, round after round, then writes SPEC.md.

Read the file before you move on. Anthropic describes the most useful specs as self-contained: they name the files and interfaces involved, state what is out of scope, and end with an end-to-end verification step that proves the feature works. If yours lacks that last part, add it now. Every later step leans on it.

Step 2: Turn the spec into a plan in plan mode

Start a new session. Anthropic's advice after the interview is to begin a fresh session, so the context is clean and the written spec is the reference. Open it in plan mode:

claude --permission-mode plan

In plan mode Claude reads files and runs shell commands to explore, but does not edit your source. Inside an existing session, Shift+Tab or a /plan prefix does the same. Now ask for the plan. This prompt is ours, modelled on the plan step in Anthropic's docs:

Read @SPEC.md and the code it touches. Create a plan: which files change,
in what order, and a numbered task list where every task names the check
that proves it. Make the first step of the plan saving it as PLAN.md and
the task list as tasks.md, then stop.

Expected result: Claude presents a plan and asks how to proceed. The options are:

  • Yes, and use auto mode: approve and let Claude run with background safety checks. It reads "Yes, auto-accept edits" where auto mode is unavailable.
  • Yes, manually approve edits: approve and review each edit.
  • No, keep planning: stay in plan mode and say what to change.

Press Ctrl+G to open the plan in your text editor and change it directly. Use "No, keep planning" until the plan matches the spec, then "Yes, manually approve edits" so you watch PLAN.md and tasks.md being written and nothing else.

Step 3: Review the plan with a subagent

The session that wrote a plan is a poor judge of it. A subagent runs in its own context window with its own tools, so it reads the files without the conversation that produced them. Save this as .claude/agents/plan-reviewer.md:

---
name: plan-reviewer
description: Checks PLAN.md and tasks.md against SPEC.md before implementation
tools: Read, Glob, Grep
model: sonnet
---

You review implementation plans. Compare PLAN.md and tasks.md with SPEC.md.
Report requirements that no task covers, tasks that name no check, and work
the spec lists as out of scope. Report gaps, not style preferences.

The frontmatter fields are the ones in Anthropic's code-reviewer example; only name and description are required, and the prompt text is ours. The tool list is read-only on purpose. Ask for the reviewer by name:

Use the plan-reviewer subagent to check PLAN.md and tasks.md against SPEC.md.

Expected result: a short list of gaps comes back to your session. Fix them in the plan files, not in your head, while changing direction still costs one edit.

For a large feature you can also split the research before planning. The subagent docs give the pattern: "Research the authentication, database, and API modules in parallel using separate subagents." It works when the areas do not depend on each other, and each subagent spends its own tokens, so save it for specs that span several parts of the codebase.

Step 4: Add hooks as quality gates

Instructions in CLAUDE.md are advisory; hooks are deterministic. That is Anthropic's distinction and the point of this step: the gates below run every time, whether or not the model remembers the rule.

GateHook eventCan it block?
Freeze the specPreToolUse, matcher Edit|WriteYes. Exit code 2 blocks the tool call and Claude is told why
No ending a turn on failing testsStopYes. Exit code 2 prevents Claude from stopping and the conversation continues
No closing a task on failing testsTaskCompletedYes. Exit code 2 keeps the task open; needs the task-tracking tools
Format every editPostToolUse, matcher Edit|WriteNo. The tool has already run; the hook only tidies the result

Start with the first two. Save the test gate as .claude/hooks/test-gate.sh and make it executable with chmod +x:

#!/bin/bash
# .claude/hooks/test-gate.sh
INPUT=$(cat)
if [ "$(echo "$INPUT" | jq -r '.stop_hook_active')" = "true" ]; then
  exit 0  # Allow Claude to stop
fi

# Run the test suite
if ! npm test 2>&1; then
  echo "Tests not passing. Fix failing tests before finishing." >&2
  exit 2
fi

exit 0

The script is assembled from two snippets in Anthropic's docs: the stop_hook_active guard from the hooks guide and the run-tests-then-exit-2 block from the TaskCompleted example in the hooks reference. Swap npm test for your own command. Then register both gates in .claude/settings.json:

{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Edit|Write",
        "hooks": [
          {
            "type": "command",
            "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/protect-files.sh"
          }
        ]
      }
    ],
    "Stop": [
      {
        "hooks": [
          {
            "type": "command",
            "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/test-gate.sh"
          }
        ]
      }
    ]
  }
}

protect-files.sh is the script from the hooks guide, copied as written except for its pattern line, which gains the spec:

PROTECTED_PATTERNS=("SPEC.md" ".env" "package-lock.json" ".git/")

Expected result: /hooks lists a PreToolUse hook and a Stop hook. Ask Claude to add a line to SPEC.md and the edit is blocked before it runs, with the script's "Blocked:" message passed back to Claude as feedback.

Three details decide how strict the test gate is:

  • The guard makes it one forced retry. stop_hook_active is true when Claude is already continuing because of a Stop hook, so with the guard a failing suite sends Claude back once per stop. Without it, the gate holds until tests pass, and Claude Code's own cap applies: it overrides the hook after eight consecutive blocks with no tool call in between.
  • Stop fires on every finished response. The hooks guide says Stop hooks run whenever Claude finishes responding, not only at task completion, so keep the gated command fast.
  • TaskCompleted is a per-task gate with a catch. It fires when a task is marked complete through the task-tracking tools. Since v2.1.268 those tools are on by default only for older models; on current ones you opt in by starting Claude Code with CLAUDE_CODE_ENABLE_TODO_TOOLS=1.

Our Claude Code hooks examples has more gates in the same format, including two stop gates. If you would rather not write configuration yet, /goal is the built-in shortcut: the hooks reference describes it as a session-scoped, prompt-based Stop hook, so /goal all tests pass keeps Claude working toward that condition for one session.

Step 5: Execute the tasks one at a time

Implementation is now the dull part, which is the goal. One task, one clean context, one check, one commit.

One task, one loop
  1. 1
    Clear the context

    Run /clear or open a new session, so the window holds the spec, the plan and nothing left over from the last task.

  2. 2
    Ask for one task

    Name it: implement task 1 from @tasks.md, following @PLAN.md. Ask Claude to run the task's check and show the output.

  3. 3
    Let the gates work

    The Stop hook runs your tests when Claude tries to finish. On a failure Claude receives the reason and keeps going.

  4. 4
    Read the evidence

    Look at the test output and the diff, not the summary. Anthropic's advice is to have Claude show evidence rather than assert success.

  5. 5
    Commit

    One commit per task, so a bad task is one revert away.

  6. 6
    Repeat

    Tick the task off in tasks.md and start the loop again for the next one.

Expected result per task: a diff that matches the task, test output you can read, and one commit. If you correct Claude twice on the same task and it is still wrong, Anthropic's guidance is to run /clear and restart with a better prompt. With the spec and plan on disk, a restart costs almost nothing.

A spec workflow gets one feature right. Our AI SaaS Builder program covers the rest of the path: idea validation, Supabase, a Claude API feature, Claude Code and MCP servers, Vercel deployment, launch and Stripe billing.

Step 6: Review the diff against the plan

Before you open a pull request, give the result to a reviewer that did not write it. This is the review prompt from Anthropic's best-practices page; swap "rate limiter" for your feature:

Use a subagent to review the rate limiter diff against PLAN.md. Check that
every requirement is implemented, the listed edge cases have tests, and
nothing outside the task's scope changed. Report gaps, not style preferences.

Expected result: the implementing session receives the gaps directly and can fix them and ask for a second review. For plain correctness bugs, the bundled /code-review skill reviews the current diff. When both come back clean, ask Claude to commit and open the pull request.

Claude Code and Spec Kit: when should you install GitHub's toolkit instead?

Spec Kit is GitHub's packaged version of this workflow, published at github.com/github/spec-kit. Its README lists Python 3.11 or later and uv as prerequisites, and two terminal commands set it up for Claude Code:

uv tool install specify-cli
specify init my-project --integration claude

The integrations reference describes the Claude Code integration as skills-based, installed in .claude/skills. The steps run inside Claude Code, not in the terminal, one at a time with a review between each. Spelling varies by agent, so type /speckit and use the names your install shows; the quickstart writes them with a hyphen.

SkillWhat it doesPath
/speckit-constitutionSets project principles, once per projectFull
/speckit-specifyCaptures what the feature is and why, in spec.mdShort and full
/speckit-clarifyAsks targeted questions and writes the answers back into the specFull
/speckit-planSets the technical approach in plan.mdShort and full
/speckit-checklistGenerates a requirements-quality checklistFull
/speckit-tasksBreaks the plan into tasks.mdShort and full
/speckit-analyzeRead-only check for conflicts and gaps across spec, plan and tasksFull
/speckit-implementImplements against those filesShort and full
/speckit-convergeChecks the code against spec, plan and tasks, and adds tasks for any gapsShort and full

The short path is specify, plan, tasks, implement and converge; the full path adds constitution, clarify, checklist and analyze. The hooks from Step 4 are independent of Spec Kit, since they fire on tool calls and stops whatever skill is running. Change the protected pattern to spec.md, because the match in the docs' script is case-sensitive.

Build your own skills or install Spec Kit?
Build your own if
  • You want two or three commands, not nine
  • Your team already has a spec format it likes
  • You want to read and edit every prompt in the workflow
  • You would rather not add a Python tool to the project setup
Install Spec Kit if
  • You want a ready-made sequence from constitution to converge
  • Several people need to run the same steps the same way
  • You want the clarify, checklist and analyze gates built in
  • You move between Claude Code and other coding agents

Does every change need spec-driven development with Claude Code?

No. Anthropic says plan mode adds overhead, and a spec adds more. Two questions sort most changes: how many files the change touches, and whether you already know the approach.

How much process a change deserves
Approach is unclear
Plan mode only: explore, read the plan, skip the files
Full workflow: interview, plan review and gated tasks
Approach is clear
Just ask. If the diff fits in one sentence, skip the plan
A short spec and gated tasks; skip the long interview
Small change
Multi-file change

If you are building without a programming background, the spec matters more, not less, because it is the part you can read and judge. Our vibe coding guide covers that way of working, and the MCP guide explains how to give Claude the issue tracker or docs a spec refers to.

Troubleshooting: where the workflow breaks

  • The spec is vague. It has no verification step. Add the end-to-end check Anthropic recommends and rerun the plan.
  • /hooks shows nothing. The hooks guide lists the causes: invalid JSON (no trailing commas or comments), a settings file in the wrong place, or a file watcher that missed the edit. Restart the session to force a reload.
  • The hook script never runs. Make it executable with chmod +x.
  • Claude cannot finish a turn. Your test command fails for a reason Claude cannot fix, such as a missing service. Keep the guard in the script, or fix the environment first.
  • The task gate never fires. TaskCompleted needs the task-tracking tools, which newer models do not load by default. Use the Stop gate, or opt in.
  • The review returns ten findings. Sort them by whether they affect the spec's requirements. Fix those and leave the style notes.

Spec-driven development with Claude Code: FAQ

What is spec-driven development with Claude Code?

It is a workflow where Claude Code writes and follows documents before it writes code. A spec states what the feature must do and how to verify it, a plan states how to build it, and a task list breaks the plan into units with a check each. Claude Code supplies the parts: skills for slash commands, plan mode, subagents for review, and hooks that keep a turn from ending while a check fails.

Does Claude Code have a built-in spec command?

No. Anthropic's command reference lists /plan, /goal, /code-review and /verify, but no built-in /spec, checked October 2026. You create one as a skill: a SKILL.md file in .claude/skills/spec/ becomes the /spec command. Anthropic's best-practices page supplies the prompt, which has Claude interview you with the AskUserQuestion tool and then write a complete spec to SPEC.md.

How do I use GitHub Spec Kit with Claude Code?

Install the CLI with uv tool install specify-cli, then run specify init with your project name and --integration claude. Spec Kit's integrations reference says the Claude Code integration is skills-based and installs its skills in .claude/skills. Inside Claude Code you run the steps in order: /speckit-specify, /speckit-plan, /speckit-tasks, /speckit-implement and /speckit-converge. Spec Kit needs Python 3.11 or later and uv.

Is plan mode the same as a spec?

No. Plan mode is a permission mode in which Claude reads files and proposes changes without editing your source until you approve the plan. A spec is a file that outlives the session and states what to build, what is out of scope and how to prove it works. Use the spec to settle what, then plan mode to settle how. Anthropic recommends a fresh session once the spec exists.

Can a hook stop Claude Code from finishing while tests fail?

Yes. A Stop hook that exits with code 2 prevents Claude from stopping, and its stderr message tells Claude why it should continue. Claude Code overrides the hook after eight consecutive blocks with no tool call in between, so a broken check cannot hold a turn open forever. A TaskCompleted hook does the same for one task, but only in sessions that have the task-tracking tools.

Do subagent reviews use extra usage?

Yes. A subagent runs in its own context window, which keeps your main conversation clean, but Anthropic's docs say each subagent spends its own tokens and its results return to the main conversation. Spend that on the two reviews that matter, the plan before any code exists and the final diff, where a fresh context is not biased toward work it just produced.

When should I skip the spec?

Use Anthropic's rule of thumb for planning: if you could describe the diff in one sentence, skip the plan and ask Claude to make the change directly. A typo, a log line or a renamed variable needs no spec. Write one when the change touches several files, when you are unsure of the approach, or when the code is unfamiliar to you. That is where the overhead pays back.

All Access · all four programs · $99/mo

Got the workflow? Build the product around it.

AI SaaS Builder, included in All Access, has a module on Claude Code and MCP servers, plus Supabase, the Claude API, Vercel deployment and Stripe integration, with the other three programs, live coaching and the private community in one subscription.

Start All Access — $99/mo →30-day money-back guarantee
Free · no signup

Start with the free Creator Starter Kit

Templates and checklists for getting a first product off the ground, free. Then read the Claude Code guide this walkthrough builds on.