Spec-driven development means writing what you are building and how you will know it works before an AI agent writes code. In practice it is three short Markdown files per feature: spec.md for what and why, plan.md for how, and tasks.md for the order of work with a check on each step. This guide shows the files, a worked example, the tools that automate the process and the ways it goes wrong.
Spec-driven development is a workflow where you write down what you are building and how you will know it works before an AI agent writes any code. In practice that is three short Markdown files per feature, kept in the repository: a spec for what and why, a plan for how, and a task list for the order of work, with a check on every step.
Tool details checked October 2026 against the GitHub Spec Kit README and its methodology document, the Kiro specs docs, the OpenSpec README and the BMAD Method README. The method described here is drawn from those sources. The worked example is illustrative, written for this guide; it is not a log of a shipped feature.
This is the method guide. It is deliberately tool-neutral: the same three files work whether you drive Claude Code, Cursor, Codex or anything else that reads a repository. For the tools themselves, our Claude Code complete guide is the place to start. This page is for solo builders and small teams who have felt an agent produce a lot of code quickly and then spent longer than that working out whether it was the right code.
What is spec-driven development?
It is a change in what you treat as the source of truth. In ordinary AI coding the prompt is disposable: you type a request, the agent builds something, and the only lasting record of what you wanted is the code it produced. In spec-driven development the request becomes a file. The Spec Kit README puts the rule in one line: define what and why before deciding how to build it.
That matters more with an agent than with a human colleague, for a plain reason. An agent fills every gap in your request with a plausible guess and does not tell you which parts were guesses. A spec moves those guesses to the front, where they are cheap to correct, instead of leaving them buried in a 600-line diff.
- "Add CSV export to the orders page"
- The agent picks the columns, limits and edge cases
- Done means it looks right in the browser
- The next session starts from the code alone
- Scope grows with every follow-up prompt
- A one-screen spec with acceptance criteria
- You decide the edge cases; open questions are marked
- Done means every criterion has a passing check
- The next session reads the spec, plan and tasks
- An out-of-scope list holds the line
None of this replaces fast, loose prompting for prototypes. Our vibe coding guide covers that style and when it is the right call. Spec-driven development is what you reach for when the code has to survive past the weekend.
The SDD workflow: spec, plan, tasks, build, verify
Every version of the method, whatever the tool, runs the same five stages. Each stage produces something you can read and reject before the next one starts.
- 01Spec
What and why. Stories, acceptance criteria, what is out of scope.
- 02Plan
How, in this codebase. Decisions, files touched, risks.
- 03Tasks
Small ordered steps, each with a check that proves it.
- 04Build
The agent implements one task at a time against the files.
- 05Verify
Run the checks, compare with the spec, update the files.
The files are small on purpose, and each answers a different question:
| File | The question it answers | What goes in it | It changes when |
|---|---|---|---|
| spec.md | What are we building, for whom, and how do we know it works? | Problem, user stories, acceptance criteria, out of scope, open questions | Requirements change |
| plan.md | How will we build it in this codebase? | Approach, decisions with reasons, files touched, risks, how to verify | The technical approach changes |
| tasks.md | In what order, and what proves each step? | Numbered tasks, each small enough to review, each with a check | Work is completed or re-sliced |
| AGENTS.md | What is always true in this project? | Stack, commands, conventions, rules that outlive any one feature | Rarely |
A layout that works for a solo builder is one folder per feature. Spec Kit and Kiro both generate something close to this; you can also just create it by hand.
specs/
export-orders-csv/
spec.md # what and why
plan.md # how
tasks.md # in what order, each with a check
AGENTS.md # standing rules that apply to every featureWhat goes in spec.md?
A spec describes behaviour, not implementation. If a sentence names a library, a table or a file, it belongs in the plan. The test of a good spec is that someone who has never seen your codebase could read it and tell you whether the finished feature is correct.
Here is an illustrative spec for a small feature in a store dashboard. The acceptance criteria use the WHEN and SHALL pattern that Kiro's docs call EARS notation, because it forces each requirement into a form you can test.
# Spec: Export orders to CSV
## Problem
Shop owners copy orders into spreadsheets by hand for their accountant.
They need a file they can download themselves.
## Users and stories
- As a shop owner, I can export my orders for a date range as a CSV file.
- As a shop owner, I only ever see my own shop's orders in the export.
## Acceptance criteria
- WHEN an owner picks a date range and clicks Export,
THE SYSTEM SHALL download a CSV with one row per order in that range.
- WHEN the range contains no orders,
THE SYSTEM SHALL show "No orders in this range" and download nothing.
- WHEN the range contains more than 10,000 orders,
THE SYSTEM SHALL email a download link instead of blocking the page.
- THE SYSTEM SHALL never include orders from another shop.
## Out of scope
- Excel (.xlsx) or PDF formats
- Scheduled or recurring exports
- Exporting customers or products
## Open questions
- [NEEDS CLARIFICATION: which timezone defines the start and end of a day?]Three parts of that file do most of the work. Acceptance criteria are the definition of done; every one should be checkable by a test or by clicking through the feature. Out of scope stops the agent from helpfully adding an Excel export you never asked for. Open questions are where you make the agent admit uncertainty: Spec Kit's templates require a [NEEDS CLARIFICATION] marker wherever the request was ambiguous, instead of a silent guess. Answer those before you move on.
What goes in plan.md?
The plan is where the spec meets your actual codebase. It records the technical decisions and, just as usefully, the reasons, so that a later session does not undo them. Ask the agent to draft it after reading the spec and the relevant code, then edit it yourself.
# Plan: Export orders to CSV
## Approach
A server route streams the CSV. No new dependencies.
## Decisions
- Route: GET /api/orders/export?from=&to= (auth required)
- The query is scoped by shop_id from the session, never from the request
- Stream rows in pages of 1,000 so memory stays flat
- Over 10,000 rows: queue a background job, email a signed link (24 h expiry)
- Dates are interpreted in the shop's saved timezone
## Files touched
- app/api/orders/export/route.ts (new)
- lib/orders/query.ts (add the date-range query)
- app/dashboard/orders/ExportButton.tsx (new)
## Risks
- Large shops: test with 50,000 seeded orders
- CSV injection: prefix cells that start with =, +, - or @
## How we verify
- Unit tests for the query and the CSV escaping
- One end-to-end test: an export as shop A never returns shop B's ordersNotice what the plan resolved. The open timezone question from the spec now has an answer. The security requirement in the spec, never another shop's orders, has turned into a concrete decision about where shop_id comes from. And the verification section names the tests before any exist. If you cannot say how you will verify a feature, the spec is not finished.
What goes in tasks.md?
Tasks turn the plan into steps small enough to review one at a time. The rule that makes them useful is that every task carries its own check. An agent given "build the export" decides for itself when it is done. An agent given task 4 below knows exactly what has to be true.
# Tasks: Export orders to CSV
- [ ] 1. Add the date-range query in lib/orders/query.ts
Check: unit test returns only orders inside the range, for one shop
- [ ] 2. Add a CSV serializer with cell escaping
Check: unit test covers commas, quotes, newlines and =SUM( cells
- [ ] 3. Add GET /api/orders/export (streams, auth required)
Check: 401 without a session; 200 with a CSV content type
- [ ] 4. Add the cross-shop test
Check: a shop A export contains zero shop B rows
- [ ] 5. Add ExportButton with date pickers and the empty state
Check: an empty range shows the message and no file downloads
- [ ] 6. Add the background job and email for large exports
Check: 10,001 seeded orders sends one email with a working linkOrder matters. The query and the security test come before the button, because a working button on top of a query that leaks data is worse than no button. Tick tasks off in the file as they pass. When a session ends or the context fills up, the next one reads tasks.md and knows where to resume.
Spec-driven development with AI: who writes what?
The common misreading of SDD is that you now write long documents by hand. You do not. The agent drafts all three files. Your job is the part it cannot do: deciding what you want, and refusing to approve a file you have not read.
- 1Describe the outcome
Two or three sentences on who needs what and why. No solution yet.
- 2Have the agent draft the spec
Ask it to mark every assumption as an open question instead of guessing.
- 3Answer the questions, cut the scope
Edit the acceptance criteria until each one is testable. Move extras to out of scope.
- 4Have the agent draft the plan
It reads the spec and the code. You check the decisions and the files it intends to touch.
- 5Generate tasks with checks
Reject any task that has no way to prove it is done.
- 6Build one task at a time
Fresh context per task or small group. The agent runs the check before it reports.
- 7Verify against the spec
Walk the acceptance criteria one by one. A reviewer agent can do the first pass.
- 8Update the files, then merge
If the build changed a decision, the plan and spec change with it.
Two pieces of agent tooling make the verify step cheaper, without changing the method. A read-only reviewer that compares the diff with the spec keeps the check independent of the session that wrote the code; our Claude Code subagents guide has a reviewer config you can adapt. And a hook that refuses to let the agent finish while tests fail turns "run the checks" from a request into a rule; see the test gates in our Claude Code hooks examples.
How much spec is enough?
The honest objection to SDD is that it is overhead. It is, and the answer is to size it. Kiro ships a quick spec that generates all three files in one pass with no approval gates. BMAD describes a process that sends small changes straight to build. The principle is the same everywhere: ceremony should scale with risk, not with habit.
A workable rule for a solo builder: if you could describe the diff in one sentence, skip the files. If the change touches more than a handful of files, crosses a security or billing boundary, or will take more than one session, write all three. Everything in between gets a spec and a task list, with the plan folded into the tasks.
Spec-driven development tools: what each one adds
Tools do not change the method. They add templates, commands and gates so you skip fewer steps. Checked October 2026:
| Tool | Files it produces | How you start | Good fit for |
|---|---|---|---|
| GitHub Spec Kit | spec.md, plan.md and tasks.md per feature, plus a project constitution | uv tool install specify-cli, then /speckit-specify, /speckit-plan, /speckit-tasks | A full, opinionated process that works with many coding agents |
| Kiro | requirements.md, design.md and tasks.md under .kiro/specs/ | Start a spec in the Kiro IDE, CLI or web app | Approval gates between phases and a task runner built in |
| OpenSpec | proposal.md, specs, design.md and tasks.md per change | npm install -g @fission-ai/openspec, openspec init, then /opsx:propose | Changes to an existing codebase, kept light |
| BMAD Method | Briefs, specifications and architecture documents | npx skills add bmad-code-org/BMAD-METHOD | A process that scales its depth to the size of the work |
| Plan mode in your agent | A plan in the conversation, no files unless you save it | Built into Claude Code and Cursor | Small changes where three files would be overkill |
| Plain Markdown | Whatever you write in a specs/ folder | A text editor | Learning the method before choosing a tool |
A few details worth knowing before you pick. Spec Kit is open source under the MIT licence, needs Python 3.11 or later and uv, and runs its steps as skills inside your coding agent, one at a time, with a review after each. Its current flow adds a project constitution at the start and a converge step at the end. Kiro is a full coding agent with specs built in, and lets you go requirements-first or design-first. OpenSpec is also MIT licensed, needs Node.js 20.19 or later, and archives each change's spec once it ships so the main specs stay current.
If you have not settled on the agent underneath, that choice comes first. Our Claude Code vs Cursor comparison covers the two most common options, both of which have a plan mode that handles the small-change end of the matrix above.
Where SDD fails: what happens when you skip steps
Each stage exists because skipping it produces a specific, recognisable failure. Knowing the symptoms is the fastest way to find which step you are shortcutting.
- No acceptance criteria. The agent decides what done means, reports success, and you discover the gaps in production. Symptom: features that work in the demo path only.
- No out-of-scope list. Every follow-up prompt adds a little, and the diff doubles. Symptom: pull requests you cannot review in one sitting.
- Plan before spec. You pick the technology first and the requirements bend to fit it. Symptom: a neat architecture for a problem nobody had.
- Tasks too large. One task is "build the feature", the context fills, and quality drops towards the end. Symptom: the last third of the work is where the bugs are.
- Tasks without checks. Progress is measured in code written, not behaviour proven. Symptom: a ticked list and a failing build.
- Specs that are never updated. The code moves on and the files describe a feature that no longer exists. Symptom: the next agent session follows the stale spec and reintroduces old behaviour.
- The specification marathon. Three days of documents before a line of code. Symptom: you have rebuilt waterfall. Cut the feature in half and ship the first half.
An SDD workflow you can start this week
You do not need to adopt a framework to find out whether this suits you. Run the method by hand on one real feature first.
- Pick one feature that will take more than a single session
- Create specs/<feature>/ with empty spec.md, plan.md and tasks.md
- Draft the spec with your agent; make it mark open questions
- Rewrite the acceptance criteria until each one is testable
- Add at least three items to the out-of-scope list
- Approve the plan only after checking the files it will touch
- Reject any task that has no check
- Build task by task, clearing context between unrelated tasks
- Walk the acceptance criteria before you merge
- Note which step you wanted to skip; that is the one a tool should enforce
A spec makes one feature go well. Stringing dozens of them into a product that people pay for is a longer discipline, and it is what our AI SaaS Builder program is built around: from validating the idea to shipping on Supabase and Next.js with an agent doing the typing.
Spec-driven development: FAQ
What is spec-driven development?
Spec-driven development, or SDD, is a way of building software where a written specification comes before the code and stays the source of truth. You describe what the feature must do and how you will know it works, turn that into a technical plan and a task list, and only then let a developer or an AI coding agent implement it against those files.
Is spec-driven development just waterfall with a new name?
No, though it can turn into waterfall if you let it. Waterfall writes one large specification up front and builds for months. SDD as practised with AI agents works one feature at a time: a short spec, a plan and tasks for a change you will ship this week, updated as you learn. OpenSpec's own description of the approach is "iterative not waterfall".
What is the difference between spec-driven development and vibe coding?
Vibe coding starts from a prompt and steers by reacting to what the agent produces. Spec-driven development starts from a written definition of done and checks the result against it. Vibe coding is faster for prototypes and throwaway work. SDD is slower to start and pays back when the code has to be maintained, extended or handed to another agent session.
Do I need a tool like Spec Kit or Kiro to do SDD?
No. The method is three Markdown files in your repository and the discipline to write them in order. Tools such as GitHub Spec Kit, Kiro, OpenSpec and BMAD Method add templates, commands and review gates around the same idea, which helps on larger features and teams. Start with plain files, and add a tool once you know which step you keep skipping.
How long should a spec be?
Short enough that you will read all of it before approving, which for most features is one screen: a problem statement, a few user stories, four to eight acceptance criteria, an out-of-scope list and any open questions. If a spec runs to several pages, the feature is usually more than one feature and should be split before any code is written.
Does spec-driven development work on an existing codebase?
Yes. You do not need to document the whole system first. Write a spec only for the change you are about to make, and let the plan reference the files it touches. OpenSpec describes itself as built for brownfield projects, Spec Kit publishes an existing-project guide, and Kiro can generate bugfix specs as well as feature specs.
How is SDD different from test-driven development?
They work at different levels and fit together. Test-driven development writes a failing test before each small piece of code. Spec-driven development writes the feature's intended behaviour before any tests or code exist. In practice the acceptance criteria in the spec become the tests, so a good spec makes test-first work easier for an AI agent to carry out.
One good spec ships a feature. A system of them ships a product.
AI SaaS Builder, included in All Access, runs this workflow across a whole build: validating the idea, planning the schema, and shipping with Claude Code and Cursor on Supabase and Next.js, with the other three programs, live coaching and the private community in one subscription.
Not ready for a program?
Start with the free Creator Starter Kit, then come back and run the checklist above on your next feature.