BMAD-METHODv6.10.0
[A] appears at the end of each micro-file.--party=anti-concensus-club), optional persistent per-party memory (memlog), and pacing/dynamics rules. Open-ended by default — ends only on an explicit signal, with an opt-in --non-interactive flag; standing agents are kept alive and resumed. 4 run modes: session / auto / subagent / agent-team (point-to-point relay, not broadcast). Ships with a preloaded "Code Review Crew" party.{memory_dir}/<party_id>/.bmad-spec or bmad-quick-dev.brain-selector.html, 108 techniques), session orchestration, clustering and prioritization. 3 selectable stances — Facilitator / Creative Partner / Ideate-for-me. Append-only memlog with optional --by user/--by coach authorship attribution.SPEC.md, the five-field kernel (Why, Capabilities, Constraints, Non-goals, Success signal) plus companion files for load-bearing content that doesn't fit the kernel. The canonical machine contract every downstream BMad skill consumes._bmad/custom/ (team-committed or user-gitignored), then verifies it resolves.@kayvan/markdown-tree-parser — generates an index.md and offers to delete, archive, or keep the original.index.md that references every file in a target folder, grouped by type/subdirectory with content-derived (not filename-derived) descriptions..memlog.md audit trail, addendum.md overflow doc, Essential Spine + Adapt-In Menu, and a Reviewer Gate with HTML report. Supersedes the deprecated create/edit/validate PRD shims. Scale-adaptive depth Level 0–4 × domain complexity Low/Medium/High.ARCHITECTURE-SPINE.md as source of truth, with a breadth-coverage rubric and a Reviewer Gate. Supersedes the deprecated fixed 11-step bmad-create-architecture flow.sprint-status.yaml as the single source of truth.bmad-review-edge-case-hunter.data-testid; each test must run independently and must not depend on order.DP.A free, open-source design workflow giving designers expert AI agents to guide them through strategy and design process — making impactful deliveries for AI-driven or traditional team development.
WDS unites entrepreneurs, developers, and designers under a shared language for the AI era. Under the hood: a structured, AI-assisted pipeline taking a product from raw idea to production-ready implementation, coordinated by three specialized AI agents.
BMAD-native: runs via the BMAD Method Module — compatible with Claude Code, Cursor, Windsurf, Codex, and 16+ IDEs. MIT license · free · community-maintained.
A method that unites
WDS is an AI Agent framework that creates a unified language for entrepreneurs, developers, and designers to collaborate in the AI era. WDS assists you in strategic design leadership for any digital product from idea to polished product.
You map user psychology and business goals. You envision the user interaction and create conceptual specifications that guide development. You transform vision into strategic design thinking. AI amplifies your expertise. With WDS, you're the strategic center that holds it all together.
WDS is a module for designers within the BMad Method. Instead of mindless prompting, the BMad method is most effective when run in an IDE like Cursor, VS Code, Windsurf, or similar tools.
The method preserves AI capabilities by dividing creative conversations into dialogues that result in high-level documents. These documents become the perfect prompts for the next step — preserving the AI's context window.
Everything is saved and published using GitHub — the same technology developers use. By delivering design in the perfect form for development, designers can deliver far more value than ever before.
Napkin sketches, whiteboard drawings, Figma prototypes, code snippets, inspiration boards — all of it. This is where creativity happens. The playful creative process reaches a point when experimentation leads to something great, something we want to build and bring to the world.
When you gather all your experiments in one place, defining them as user scenarios with step-by-step interaction and clear explanations — not just what you designed, but why — you give AI all it needs to make the final product reality.
In the world of AI development, the specifications become the product. The code is just output — generated and regenerated from your design work. Your specifications are the source of truth that gets maintained and refined over time.
Want to change something? Update the specifications. AI regenerates the code — clean, consistent, and aligned with your design intent. No spaghetti code. No lost reasoning. No painful refactoring. Your design work IS the product.
"Whiteport Design Studio adds a strategic design dimension that strengthens BMad: a practical way to translate intent into interfaces and product direction. It fits the spirit of the agent framework, with expert-driven workflows where AI is a partner, not an autopilot."
Saga drives strategy in analysis (alignment · project brief · trigger mapping). Freya drives design in planning — PRD → Phase 3 scenarios → Phase 4 (UX design · conceptual specs · spec audit) → Phase 6 assets → Phase 7 design system (optional). After Phase 7, two workflows run in parallel: Design Delivery (Phase 4, Freya — packages flows for development handoff) and Create UI Prototype (Phase 5, Mimir — validates design in code). Mimir then drives the build in implementation (tech audit → PRD → build, one verified requirement at a time), reading the WDS specs — governed by BMAD Dev Story discipline and gated by Murat (TEA). Freya's Phase 8 runs as standalone brownfield support.
Mimir owns planning (tech-audit → write-prd, derived straight from the WDS page specs so design intent is preserved — something a plain BMAD Dev Story can't read). BMAD Dev Story then covers execution for anything outside Mimir's build loop: TDD implementation → systematic debugging → code review, run by Amelia. Dev Story's own task planning is intentionally deferred to Mimir for WDS-enabled stories — write-prd is the single planning authority, avoiding two systems owning the task list. The build/story overlap is not a conflict: Mimir's "one requirement at a time" is the rule, and Dev Story's TDD discipline is the mechanism that realizes it.
Test Architect Enterprisev1.19.0
Separate Installation
BMAD TEA (Test Architect Enterprise) is a standalone add-on module installed on top of BMAD core to extend comprehensive test architecture capabilities. While BMAD core provides bmad-qa-generate-e2e-tests for quick test needs within a sprint, TEA elevates to a different strategic tier — from risk assessment to CI/CD governance, from contract testing to NFR evidence audit.
TEA does not replace BMAD core — TEA complements core. All TEA workflows integrate with BMAD story/PRD/epics artifacts and use the same config pattern (TOML customization layer).
| Comparison Dimension | BMAD Core | BMAD TEA |
|---|---|---|
| Installation | Included | Separate add-on |
| Test agent | Amelia (QA role) | Murat 🧪 (specialized) |
| Test strategy | Happy path + critical errors | Risk-based, full coverage matrix |
| Test levels | API + E2E | Unit + Component + API + E2E + Contract + NFR |
| CI/CD | Run test command | CI pipeline scaffold + quality gates |
| Knowledge base | — | 40+ fragments, JIT loading |
| Release gate | — | Trace coverage + NFR evidence audit |
or "Test Architect"
"Strong opinions, weakly held" — Murat's mantra. Always willing to change perspective when new data emerges.
Speaks in risk calculations and impact assessments. Combines data with intuition — no purely theoretical judgments, always grounded in project reality.
Playwright · Cypress · pytest · JUnit · Go test · xUnit · RSpec · k6
GitHub Actions · GitLab CI · Jenkins · Azure DevOps · Harness CI
| Code | Skill | Description |
|---|---|---|
| TMT | bmad-teach-me-testing | Interactive testing academy — 7 sessions from fundamentals to advanced, with quizzes and role-based learning paths. |
| TD | bmad-testarch-test-design | Design test strategy — risk assessment, NFR planning, coverage strategy for a system or epic. |
| TF | bmad-testarch-framework | Scaffold a production-ready test framework — Playwright or Cypress, with fixtures, helpers, and full configuration. |
| CI | bmad-testarch-ci | Advise and scaffold a CI/CD quality pipeline — quality gates, sharding, parallelization, reporting. |
| AT | bmad-testarch-atdd | ATDD — generate failing acceptance tests + implementation checklist before dev writes any code. |
| TA | bmad-testarch-automate | Expand coverage — generate API/E2E tests, fixtures, DoD summary for a story or feature. Parallel subagents. |
| RV | bmad-testarch-test-review | Review quality of written tests — audit against the knowledge base and best practices, deliver specific improvements. |
| NR | bmad-testarch-nfr | NFR Evidence Audit — assess implemented NFR evidence, recommend missing actions. |
| TR | bmad-testarch-trace | Trace coverage & Release Gate — map requirements → tests (5 steps), final step yields a decision: PASS / CONCERNS / FAIL / WAIVED. |
| GATE | routing intent NEW v6.10.0 | Composite release-gate shortcut — walks optional test-review → optional nfr-assess → trace Phase 2 gate, without merging the underlying workflows. |
If the intent is clear (e.g., "hey Murat, design tests for this epic"), Murat skips the menu and dispatches directly to the appropriate skill.
The 🧪 prefix in every response indicates Murat is active.
- 01Assess & ProfileGather role, experience level, and learning goals. Check for an existing progress file.what
Collect role (QA / Dev / Lead / VP), experience level (Beginner / Intermediate / Experienced), learning goals, and optional pain points. Detect existing progress file.
whyDepth and examples adapt per role — a QA needs practice focus, a VP needs strategy and metrics context.
howLoad or create
{user}-tea-progress.yaml; pre-select recommended session path based on experience. - 02Select SessionChoose from the 7-session hub menu; return here between sessions.what
Present 7 sessions with status (✅ completed / 🔄 in-progress / ⬜ not-started), scores, and duration. Allow non-linear selection.
whyNon-linear navigation lets experienced learners jump directly to advanced content without re-covering basics.
howSelect 1–7 to enter a session; X exits and saves progress; all-complete routes automatically to certificate generation.
- 03TeachDeliver role-adapted content for the chosen session using TEA knowledge fragments.what
Present concepts, examples, and TEA knowledge fragments adapted to learner role. Sessions cover: Quick Start → Core Concepts → Architecture → Test Design → ATDD & Automate → Quality & Trace → Advanced Patterns.
whyJust-in-time content loading keeps each session focused and context-window manageable across the full 24k-line TEA knowledge base.
howLoad session-content-map for this session; select role-appropriate examples; reference local docs and knowledge fragments; web-browsing as fallback for updated framework docs.
- 04Quiz & ValidateKnowledge check after each session — ≥70% to advance.what
3 targeted questions per session with constructive, role-aware feedback. Up to 3 attempts; score saved to progress file.
whyActive recall cements learning and surfaces gaps before the learner moves on to the next session.
howScore ≥70% proceeds; <70% offers focused review of the specific topics that failed. Session 7 (Advanced Patterns) is exploratory — auto-scores 100% on completion.
- 05Generate ArtifactsSave session notes, update progress file; issue completion certificate when all sessions are done.what
Write
session-notes.mdfrom template; update{user}-tea-progress.yaml(status, score, completion %). Generate completion certificate when all 7 sessions reach ≥70%.whyLearners keep structured notes and a persistent progress record that survives across days and conversations.
howAuto-save after each quiz completion. Certificate includes all session scores, skills acquired, and next-step recommendations.
brain-selector.html, 108 techniques), and an append-only memlog with optional --by user/--by coach authorship attribution.- 01Select techniqueMind-map, SCAMPER, reverse thinking — choose the method that fits the context.what
AI initializes the session with a 3-question interactive setup then presents a technique selection menu.
- 1AI asks 3 setup questions to define session scope
- topic — main topic of the brainstorm
- goal — broad exploration or focused ideation (deep)
- capture session — whether to save the output document (default Yes)
- 2AI offers 4 technique selection modes
- Mode 1 — User-selected: user browses library of ~60 techniques and picks
- Mode 2 — AI-recommended: AI analyzes context, suggests 3-4 best fits with one-line rationale
- Mode 3 — Random: AI picks randomly (good for breaking patterns)
- Mode 4 — Progressive flow: moves sequentially from broad → narrow through multiple techniques
- 3User confirms choice before moving to step 02 — no auto-progress
why- ·No suitable technique → session becomes unfocused or falls into anchor bias
- ·Each technique opens a different angle of attack on the problem space
- ·SCAMPER is effective when improving existing solutions (Substitute / Combine / Adapt / Modify / Put to other use / Eliminate / Reverse)
- ·Reverse Brainstorming is effective when stuck — ask "how to create the problem?" instead of "how to solve it?"
- ·Six Thinking Hats is effective when balancing multiple perspectives is needed (facts / emotions / risks / benefits / creative / process)
- ·Mind Mapping is effective for broad exploration without a specific scope
howAI loads the technique library (~60 entries) from the internal knowledge base then routes by mode.
- 1User-selected mode: AI displays the full list with category labels (divergent / convergent / lateral / structured)
- 2AI-recommended mode: AI analyzes 2 axes
- user context (industry, problem type, stage)
- strengths of each technique (internal matching matrix)
- returns top 3-4 techniques with a one-line justification
- 3After user picks, AI confirms with a summary: "OK, we'll use [technique] for [goal]. Ready to start?"
- 1
- 02Generate ideas in batchesTarget 100+ ideas, quantity-first principle, no judgment during generation.what
AI opens the generation session, prompts according to the chosen technique, generates ideas in batches of 10-15 ideas/round, targeting 100+ ideas total.
- 1AI prompts the user according to the technique template (e.g., SCAMPER: "What can we Substitute in the current design?")
- 2Each round is timeboxed ~5-10 minutes, AI applies 3 principles
defer judgment— no evaluation of any idea during this roundbuild on others— combine existing ideas into new oneswild ideas welcome— outrageous ideas are encouraged
- 3After each round, AI asks for brief feedback ("good direction?") then continues
- 4AI does not self-generate ideas but pulls ideas out of the user via probing questions
- 5Track count continuously, target 100+ ideas before moving to convergence
why- ·Per creativity research: the first 20 ideas are usually obvious, ideas 50-100 are the "golden" zone
- ·Early critique kills divergent thinking — weak ideas still have value as stepping stones to good ones
- ·Quantity-first breaks self-censorship — users verbalize minor ideas more easily when they know filtering comes later
- ·Timeboxed rounds maintain energy — open-ended sessions tend to degrade in quality after 30 minutes
how- ·AI uses probing questions such as: "What if cost was zero?", "What would Tesla do?", "What's the opposite approach?"
- ·AI uses analogical prompting: "What if this were like X industry?" to break domain anchoring
- ·Each idea is recorded with a 1-line description, no expansion required — speed > polish at this stage
- ·AI counts silently and signals when milestones are reached ("50 ideas reached, keep going")
- 1
- 03Anti-bias pivotEvery 10 ideas, shift perspective / persona to break cognitive bias.what
After every ~10 ideas, AI actively shifts the lens to break patterns and open new idea categories.
- 1AI tracks session metadata: idea count, which technique, which cluster is dominating
- 2When threshold is reached (every 10/20/30 ideas), AI self-pivots using one of 3 methods
- Persona shift: "Now think as a CFO / kid / hacker / regulator"
- Hat shift: switch Six Thinking Hat (white = facts → red = emotions → black = risks → green = creative...)
- Question inversion: "How would we make this feature FAIL?"
- 3AI does not ask permission but declares "pivot now — let's look from X angle" and continues
- 4After a pivot, new-round ideas are often more outlier → flag separately to avoid early merging
why- ·Anchoring bias: the brain anchors to the initial framing — pivots break that anchor
- ·Availability heuristic: the next idea tends to be close to the last one → homogeneous clusters
- ·Functional fixedness: only seeing 1 use case per thing — persona shifts open alternative use cases
- ·Active pivots (without waiting for user request) create structured serendipity
how- ·AI uses an internal perspective rotation matrix: 6 personas × 3 timeframes × 4 constraint levels
- ·Persona shifts in a 60-minute session: 3-4 different personas, ~15-20 ideas per persona
- ·Standard Six Hats sequence: white (data) → red (feelings) → black (risk) → yellow (benefit) → green (creative) → blue (process)
- ·Question inversion technique from Reverse Brainstorming: list ways to cause the problem, then negate them
- 1
- 04Cluster ideasGroup ideas by theme, name clusters, remove duplicates.what
Convergence mode begins. AI groups ideas into clusters by affinity, names categories, and deduplicates.
- 1AI scans all 100+ ideas and proposes 5-9 initial clusters
- 2Cluster names use "categories everyone understands"
- e.g., UX improvements, Revenue models, Tech debt, Compliance
- avoid jargon or team-internal names
- 3User approves or renames each cluster
- 4AI assigns each idea to exactly 1 cluster
- 5Outlier ideas (one-of-a-kind, don't fit any cluster)
- placed in a separate parking lot
- never forced into a cluster (preserves originality)
- 6Duplicate ideas merged by content (merge count → signals cluster importance)
why- ·100+ spread-out ideas are noise — not actionable, not communicable
- ·Clustering turns noise into signal — reveals themes the session naturally gravitated toward
- ·"Populous" themes are usually the most investment-worthy areas of the problem space (revealed preference)
- ·"Sparse" themes also carry information — gaps may be blind spots or less important areas
- ·Per Miller's Law (7±2): human short-term memory holds ~5-9 items — clusters in this range are easier to retain
how- ·AI uses affinity mapping protocol: physical metaphor (sticky notes on a wall) even in digital form
- ·Clustering is conceptual — by purpose/audience/mechanism, not by similar wording
- ·After clustering, count ideas/cluster → bar chart visual so the user can see distribution
- ·Parking lot outlier ideas are often material for a future epic — do not discard
- 1
- 05PrioritizeAssess impact vs. effort, select top ideas to develop further.what
AI applies a prioritization framework, scores ideas, and selects the top 5-10 to carry into Phase 2.
- 1AI offers 2 prioritization frameworks
- Impact/Effort 2×2 matrix: quick wins / big bets / fill-ins / thankless tasks
- MoSCoW: Must have / Should have / Could have / Won't have
- 2User selects framework (default: Impact/Effort — more visual)
- 3Scoring process
- Impact (1-5): "if shipped, how much user/business value does this deliver?"
- Effort (team-weeks): "how long to build and ship?"
- AI proposes initial scores, user adjusts
- 4Plot ideas on the visual matrix
- 5Output: top 5-10 ideas with short rationale (1-2 lines each)
- 6Quick wins (high impact + low effort) flagged separately for immediate action
why- ·A great brainstorm without prioritization = a dead document
- ·Prioritization forces explicit trade-offs and prevents cherry-picking later
- ·Scoring surfaces hidden disagreement: if 2 people score very differently, different assumptions need to be clarified
- ·Quick wins shipped early create momentum + early evidence for the Phase 2 PRD
- ·Big bets need additional research/validation before commitment — flag to avoid hasty decisions
how- ·AI uses Impact/Effort 2×2 in Eisenhower box style — 4 quadrants with pre-defined names
- ·Effort dimension: AI estimates using reference class forecasting — compares against similar features with known history
- ·Impact dimension: AI tries to quantify ("saves 2 hours/user/week") rather than using generic ratings ("high")
- ·Avoid vanity metrics: prioritize retention/engagement over page views/downloads
- 1
- 06Save reportExport brainstorming-report.md, recommend next workflow.what
AI generates
brainstorming-report.mdcontaining the full session and saves it to_bmad-output/planning-artifacts/.- 1Report sections in order
- Setup: topic, goal, technique(s) used, date, attendees
- Full idea list: numbered, grouped by round
- Clusters: with count + name + member ideas
- Prioritization matrix: visual 2×2 or MoSCoW table
- Top picks: 5-10 ideas with rationale
- Next steps: recommended workflow to continue
- 2Frontmatter attaches metadata:
stepsCompleted,workflowType: brainstorming,date - 3AI invokes
bmad-helpwith arg = "brainstorming" - 4bmad-help suggests next workflow based on project state
- typically: research (validate ideas) or product-brief (document vision)
- if ideas are already solid: go straight to create-prd
why- ·Output artifact = single source of truth for subsequent phases
- ·Without a file, context is lost when changing chat sessions (core BMAD principle: "artifacts drive state")
- ·Teammates cannot re-derive results from a verbal summary alone
- ·Frontmatter metadata enables workflow continuation — re-invoking a workflow knows which steps are already done
how- ·Append-only writing: AI writes section by section, does not regenerate the file
- ·File path resolves via config:
{planning_artifacts}/brainstorming-{slug}-{date}.md - ·
bmad-helpauto-invokes at the final step of every workflow — uniform design pattern - ·Roadmapping syntax: next phase, next workflow, prerequisite, estimated time
- 1
- 01Define research questionsClarify the specific questions to answer (market size? tech feasibility? compliance?).what
AI turns a vague question into 5-8 sub-questions answerable with evidence.
- 1AI asks user to select one or more research scopes
- Market — TAM/SAM/SOM, competitors, trends, pricing
- Domain — vocabulary, expert models, regulations
- Technical — feasibility, stack options, implementation paths
- 2AI applies 5 Whys to drill from the original question down to the root question
- From each answer, AI asks 'why?' again
- Stop at level 5 or when a researchable question is reached
- 3AI converts root question into 5-8 specific, evidence-answerable sub-questions
- 4User validates the sub-question list before AI begins gathering
why- ·A vague question leads to wandering research — specificity creates clear coverage checklists
- ·Small sub-questions enable parallel investigation and a clear signal for 'done'
- ·Surfaces scope creep early — one original question can imply 10 sub-areas
how- ·Use the problem framing canvas: problem / users affected / current solutions / why insufficient
- ·Each sub-question must be answerable with evidence, not an opinion question
- ·Standard format: "Are there competitors using on-device LLM for document Q&A, and what is their pricing?"
- 1
- 02Gather sourcesWeb search, docs, interviews, industry reports — record sources completely.what
AI gathers evidence from 3+ independent source types, attaching a full citation to each finding.
- 1AI searches in parallel across 4 source types
- Web search — news, blogs, industry reports
- Official docs — specs, standards (ISO/RFCs), regulatory texts
- Expert content — podcasts, talks, academic papers
- Interview notes — if the user has existing user research
- 2Each finding is assigned a quality grade
- Primary — regulator docs, SEC filings, official statements
- Secondary — Reuters/Bloomberg, peer-reviewed research
- Tertiary — Wikipedia, aggregators, blog opinions
- 3Each citation includes: URL + publish date + author name + verbatim quote or clear paraphrase
why- ·Triangulation — a fact appearing in 3 independent sources = evidence; 1 source = assumption
- ·Mandatory citations prevent AI from hallucinating stats — must attach a specific URL
- ·Quality grading helps downstream synthesis weight evidence appropriately
how- ·Desk research protocol: primary first, then secondary, tertiary only for cross-checking
- ·Multi-query strategy: rephrase each sub-question 2-3 different ways to cover the keyword space
- ·Recent-first filter: prioritize sources from the past 12 months for fast-changing topics
- 1
- 03Analyze dataCross-reference, identify patterns, contradictions, and gaps in the information.what
AI cross-references data and looks for 3 types of signal: patterns, contradictions, gaps.
- 1AI builds a finding matrix
- Rows — sub-questions from step 01
- Columns — sources gathered
- Cells — quote/data from that source for that question; empty cell = gap
- 2AI scans the matrix for 3 signal types
- Patterns — ≥3 sources agree on X → high-confidence finding
- Contradictions — source A says X, source B says ~X → highlight and investigate
- Gaps — empty cells in the matrix → unanswered questions
- 3AI does not skip contradictions or gaps — both are predictive of risk
why- ·Pattern recognition is easy to overweight — contradictions and gaps are often overlooked but carry high value
- ·Contradictions are typically contested ground worth digging into
- ·Gaps are often unknown unknowns — predictive of risks in Phases 3-4
how- ·Matrix rendered as a table (sub-questions × sources); empty cells naturally surface as gap indicators
- ·Pattern threshold: ≥3 independent sources = pattern; 2 = weak signal; 1 = assumption
- ·Contradictions: AI presents both sides with evidence weight, does not pick a winner
- 1
- 04Synthesize findingsExtract key insights, quantify where possible, and highlight risks.what
AI synthesizes into 3-5 actionable insights, quantifying when possible.
- 1Each insight has the format
- Claim — 1 line, specific, quantified when data is available
- Evidence pointer — cite which source numbers in the matrix
- Confidence level — high / medium / low
- 2AI lists 3-5 risk flags cross-category
- Market — timing, demand, competition
- Technical — feasibility, scalability, vendor
- Regulatory — compliance, data, jurisdiction
- 3Output: 1 document readable in 5 minutes by a PM/founder
why- ·Synthesis ≠ summary — synthesis creates new understanding by combining ('because X and Y, therefore Z')
- ·Quantifying forces precision — 'growing fast' could be 5% or 50%, a critical difference
- ·Confidence level helps PM weight insights when they conflict with gut feeling
how- ·Pyramid Principle (Minto): conclusion-first, supporting evidence below
- ·Confidence calibration: high = 3+ primary sources; medium = 2 sources; low = 1 + inference
- ·Each insight tagged with evidence count (X sources) for quick reviewer judgment
- 1
- 05RecommendationsProvide specific recommendations based on collected evidence.what
AI provides evidence-based recommendations — each with action, rationale, owner, metric.
- 1Each recommendation has the format
- Given — evidence from insights
- Recommend — specific action
- Because — reasoning linking evidence → action
- Owner — PM / Tech Lead / Founder
- Success metric — how to know it worked
- Risk if skipped — consequence of inaction
- 2AI avoids opinion-only — each recommendation must trace back to evidence from steps 03/04
- 3If evidence is weak, AI flags low-confidence recommendation instead of pretending it's solid
why- ·Research without recommendations = trivia — PMs need decisions, not facts
- ·Given/Recommend/Because format makes recommendations defensible in meetings with stakeholders
- ·Risk if skipped creates a 'cost of inaction' — helps prioritize against other proposals
how- ·Argumentation theory: claim + grounds (evidence) + warrant (reasoning)
- ·Ranking: impact × confidence × urgency
- ·AI offers 2-3 alternative paths when there is genuine uncertainty instead of forcing 1 answer
- 1
- 06Save research reportExport research-findings.md with full citations and next actions.what
AI generates
research-findings.mdwith full citations in APA/Chicago format.- 1Report structure
- Executive summary — 1 paragraph
- Methodology — scope, sources, date range
- Sub-questions + evidence — 1 section/question
- Insights — 3-5 with confidence levels
- Recommendations — ranked with rationale
- Full bibliography — alphabetical, complete citations
- 2Frontmatter: dates, methodology, scope,
workflowType: research - 3AI invokes
bmad-help→ suggest next workflow based on strength of findings
why- ·Citations ensure reproducibility — PM can trace every claim back to its source
- ·Methodology section enables scope critique — reviewer can see if key sources are missing
- ·Artifacts drive state — without a file, context is lost when switching chats
how- ·Format: APA (Author, Year, Title, URL) or Chicago (with footnotes)
- ·Append-only writing — AI writes each section, does not regenerate the file
- ·End of file includes a 'Last updated' timestamp to track freshness
- 1
- 01Init & contextLoad prior artifacts (brainstorming, research), define the brief scope.what
AI loads prior artifacts and scopes the brief for the target audience.
- 1AI checks the workspace for Phase 1 artifacts
brainstorming-report.mdresearch-findings.mdprfaq-{project}.md
- 2If found: AI extracts relevant content — no re-derivation
- 3AI asks 3 scoping questions
- Audience — founder / engineering / mixed / stakeholder external
- Decision driven — what decision does this brief inform?
- Level of detail — 1-2 pages (preferred) or deeper
- 4AI propose outline → user adjust → commit
why- ·Brief without scope → becomes a PRD lite — long, rambling
- ·A proper brief must be 1-2 pages, stakeholders read it in 5 minutes for a go/no-go decision
- ·Extracting (not re-deriving) saves time and maintains consistency with earlier artifacts
how- ·Default audience: mixed (founder + early team) if user does not specify
- ·Artifact extraction: AI scans frontmatter + headings, lists relevant sections for user confirmation
- ·Outline proposed as bullet list for user to adjust before writing
- 1
- 02Vision & missionState the long-term vision and product mission concisely.what
AI co-drafts Vision (5-10 years) and Mission (1 sentence, what-for-whom-how).
- 1Vision drafting
- AI ask: 'What does world look like in 5-10 years if you succeed?'
- AI propose 2-3 draft options
- User selects + iterates until no more words are needed
- 2Mission drafting with the Geoffrey Moore template
- 'For [user], who [pain], our [product] is [category] that [value]. Unlike [alternative], we [differentiator].'
- Compress down to the final 1-2 sentences
- 3Validate: vision is distant from tactics; mission is close to tactics
why- ·Vision alone = a dream with no daily direction
- ·Mission alone = labor without purpose, losing long-term direction
- ·The two are complementary and serve as an anchor whenever feature creep occurs
how- ·Opposite test: if the opposite is not plausible → vision is too bland, redo
- ·Read aloud test: if it sounds generic / applies to any startup → re-draft
- ·Vision/Mission must be finalized before step 03 — all future decisions reference back to them
- 1
- 03Target usersSketch 1-3 primary personas with specific needs and pain points.what
AI sketches 1-3 personas: demographics, JTBD, pain, current alternative, success criteria.
- 1Each persona format
- Demographics — age, role, location, tech savviness
- JTBD — 'When [situation], I want to [motivation], so I can [outcome]'
- Primary pain — 1-2 specific lines
- Current alternative — how they solve it today (including 'do nothing')
- Success criteria — measurable outcome if pain resolved
- 2AI does not merge all users into 1 'average user'
- 3Validate with user
- 'Do you know a real person who matches this persona?'
- If not → persona may be fiction, redo
why- ·Average user is fiction — real users have different jobs/pain
- ·Products trying to serve everyone usually serve no one
- ·JTBD framework focus on motivation rather than demographic stereotype
how- ·AI uses Clayton Christensen's JTBD framework
- ·AI uses the empathy map as a supplementary tool: says / thinks / does / feels
- ·Also define the anti-persona — who we explicitly do not serve
- 1
- 04Problem statementDefine the problem being solved, why it matters and why it's urgent.what
AI writes a Problem Statement specific enough to be falsifiable.
- 14 components of the problem statement
- What — specific problem, not a solution disguised as a problem
- Who — who is affected, link to personas
- Why important — frequency × severity × visibility
- Why now — timing: regulation, tech enablement, market shift
- 2AI avoids solution-in-the-problem (e.g., 'we need a chatbot' = solution, not a problem)
- 3Statement must be falsifiable: can be refuted by evidence
why- ·Vague problem → vague solution → shipping features that don't address the core issue
- ·The problem statement is a contract between the team and reality
- ·Being specific enables measuring real success, not just 'feature shipped'
how- ·PAIN framework: frequency × severity × visibility
- ·Falsifiability test: 'What evidence would prove this problem doesn't exist?'
- ·AI uses 5 Whys to dig past surface complaints to the root problem
- 1
- 05Value propositionUnique value delivered to users, differentiated from existing solutions.what
AI defines the Value Proposition: unique value vs. alternatives, mapping pains → gains.
- 1AI uses the Value Proposition Canvas
- Customer pains ← linked to problem statement
- Customer gains — positive outcomes desired
- Pain relievers — how product addresses each pain
- Gain creators — how product creates each gain
- 2AI maps value vs. 3 alternatives: top competitors + current workaround + 'do nothing'
- 3Articulate differentiator: 1-2 'only we have because [moat]'
why- ·Most products fail not because they are bad, but because they are not different enough for users to switch
- ·'Do nothing' is always the strongest competitor — switching cost > pain = no adoption
- ·Articulating the value prop early reveals: if you cannot articulate it, it may not exist
how- ·Value Proposition Canvas by Strategyzer/Osterwalder
- ·Differentiation matrix: rows = features, cols = us + top 3 competitors
- ·AI flags features that all competitors have = table stakes, not differentiators
- 1
- 06Success metricsKPIs and OKRs to measure success — short-term and long-term.what
AI defines Success Metrics: 1-2 north star + 3-5 supporting KPIs.
- 1North Star Metric — 1-2 metrics that, if improved, improve the entire business
- Airbnb = nights booked; Spotify = time listening; Slack = messages sent
- 2Supporting KPIs per the HEART framework
- Happiness — NPS, satisfaction
- Engagement — DAU/MAU, session frequency
- Adoption — signup conversion, time-to-first-value
- Retention — D7/D30, cohort curves
- Task success — completion rate, error rate
- 3Separate timeframes
- Short-term (3-6 months) — leading indicators
- Long-term (1+ year) — lagging indicators
why- ·Metrics are a contract between the team and the board
- ·Vanity metrics (page views, downloads) are easy to grow but do not correlate with business value
- ·The North Star aligns the entire team around a single number
how- ·OKR template + HEART framework by Google
- ·Each metric: baseline (current) + target (12 weeks) + measurement source
- ·AI flags a warning if a metric has no baseline/target → 'vague metric'
- ·Layer North Star + 2-3 leading indicators — avoid putting all expectations on a single number
- 1
- 07ConstraintsBudget, timeline, resources, technical and regulatory constraints.what
AI lists Constraints: budget, timeline, team, technical, regulatory — categorized as MUST/SHOULD/MAY.
- 1AI asks the user about 5 types of constraints
- Budget — dollar amount, runway
- Timeline — launch date, milestones
- Team — size, skills, hiring plan
- Technical — must integrate with X, must use Y stack
- Regulatory — HIPAA / GDPR / SOC2 / PCI / industry-specific
- 2Each constraint categorized into 3 tiers
- MUST — deal-breaker if violated
- SHOULD — preferred
- MAY — nice-to-have
- 3AI flags conflicts between 2 MUST constraints — must be resolved before committing
why- ·Constraints define solution space — not listing them = assuming they don't exist = surprise mid-sprint
- ·A constraint can become a differentiator (e.g., 'on-device' due to privacy constraint becomes a USP)
- ·MUST constraint conflicts must be resolved before committing — not mid-development
how- ·Constraint analysis matrix: type × tier × source
- ·AI asks regulatory questions per industry: healthcare → HIPAA, finance → SOC2/PCI, EU → GDPR
- ·Brand constraints are often forgotten: voice, visual identity, tone restrictions
- 1
- 08Next stepsSave product-brief.md and recommend moving to Phase 2 (PRD).what
AI saves
product-brief.mdand clearly hands off to Phase 2.- 1Final document structure
- Executive summary — 1 paragraph (written last, placed first)
- Vision + Mission
- Target users — 1-3 personas
- Problem statement
- Value proposition + differentiation
- Success metrics — north star + KPIs
- Constraints — tier categorized
- 2Frontmatter:
workflowType: product-brief,stepsCompletedcomplete - 3AI invokes
bmad-help→ recommends next: create-prd or prfaq
why- ·The brief is the input for the PRD — the PRD source-extracts the brief automatically, no re-explaining
- ·Executive summary written last but placed first per the Pyramid Principle
- ·1-2 page format enforces tight writing — quality over quantity
how- ·Append-only writing: write each section, do not regenerate
- ·File path:
{planning_artifacts}/product-brief.md - ·
bmad-helpauto-invoked at the final step — uniform design pattern for all workflows
- 1
- 01Press ReleaseWrite a hypothetical press release for launch day — headline, subhead, customer quotes.what
AI writes a hypothetical press release for launch day in AP/Reuters style.
- 1PR structure (~1 page)
- Headline — 8-12 words, attention-grabbing
- Subhead — 1 sentence expand on headline
- Lead paragraph — who/what/when/where/why
- Body — 2-3 paragraphs detail
- Customer quotes — 2-3 concrete, specific quotes (e.g., 'saved me 4h/week')
- Leadership quote — articulate the theory of change
- Call-to-action — where to sign up / learn more
- 2AI writes as a freelance journalist — if it feels boring, the product is not exciting enough yet
why- ·Amazon's Working Backwards — Jeff Bezos's methodology
- ·If you cannot write a compelling PR, the product is not ready to build
- ·PR forces articulating the value proposition in user-facing language, not engineer-facing
- ·Customer quotes force specific outcomes — vague benefits = vague product
how- ·Inverted pyramid: most important info top, details bottom
- ·Boredom test: read aloud — if it feels cringe, redo
- ·Customer quotes format: '[Product] saved me [specific outcome]' — not 'I love it'
- 1
- 02External FAQFAQ for customers: how to use, pricing, benefits, comparison with other products.what
AI generates External FAQ (8-12 Q/A) from the perspective of a skeptical customer.
- 1FAQ covers the most uncomfortable topics
- Pricing — 'Why this expensive?', 'Free tier?'
- Comparison — 'How different from [competitor]?'
- Risk — 'Data privacy?', 'Vendor lock-in?'
- Migration — 'How to migrate from current solution?'
- Support — 'What if it breaks?', 'Refund policy?'
- 2The hardest questions at the top of the FAQ — not buried at the bottom
- 3Each answer 2-4 lines: direct + concrete, no marketing fluff
why- ·FAQs expose uncomfortable questions the team typically avoids
- ·Answering before launch = ready to defend the product; unable to answer = not yet solid
- ·Hardest questions on top — if the team is not sure = must resolve before committing
how- ·Devil's advocate prompting: AI generate FAQ as if hostile to product
- ·Answer format: direct + evidence + escape hatch if applicable
- ·Avoid marketing fluff — concrete numbers/comparisons rather than adjectives
- 1
- 03Internal FAQInternal FAQ: risks, assumptions, costs, timeline, technical trade-offs.what
AI generates the Internal FAQ for the team/board: risks, costs, assumptions, failure modes.
- 1Internal FAQ covers
- Top risks — market, tech, team, regulatory, financial
- Cost breakdown — dev cost, COGS, marketing CAC
- Timeline assumptions — '6 months — what if it takes 12?'
- Technical trade-offs — chosen path vs alternatives
- 'What if we're wrong about X' — pre-mortem questions
- 2AI uses the pre-mortem protocol
- Assume product fails in 12 months — tell the story of WHY
- Top 5 failure reasons → converted into a risk register
why- ·Internal FAQ is the pre-mortem mechanism
- ·Per Kahneman: pre-mortem improves forecast accuracy by 30%
- ·Surface hidden assumptions before committing — after committing it costs 10-100x to fix
how- ·Each risk: probability × impact + early warning signal + mitigation
- ·AI prompts hardest hypotheticals: 'What if Apple does this in 6 months?'
- ·Risk categories: Market / Technical / Team / Regulatory / Financial
- 1
- 04Customer ExperienceNarrate the detailed end-to-end experience from the user's perspective (moments of truth).what
AI narrates the Customer Experience end-to-end: awareness → habit → renewal.
- 15 stages with target emotion
- Awareness — how user hears about product
- Consideration — what convinces them to try
- Onboarding — first 5 minutes after signup
- First use — aha moment / value realization
- Habitual use — daily/weekly engagement loop
- 2AI identifies moments of truth at each transition
- Critical decision points where user can drop off
- Typically at: pricing page, first-use friction, error recovery
why- ·A static features list says 'what exists'; CX narrative says 'how it feels'
- ·UX problems mostly occur at transitions between stages, not in the stages themselves
- ·Peak-end rule (Kahneman): users remember peak emotion + end emotion → design accordingly
how- ·Service blueprint: above-the-line (user-visible) + below-the-line (backend)
- ·User journey mapping with emotion curve
- ·Deliberately create 1 peak emotion + strong end emotion
- 1
- 05ValidationReview everything, stress-test with hard questions, finalize the prfaq file.what
AI stress-tests the entire PR/FAQ from 5 adversarial perspectives, patches gaps, finalizes.
- 1AI plays adversarial reviewer from 5 angles
- Investor — 'TAM? Unit economics? Moat?'
- Customer — 'Why would I switch? Risk?'
- Engineer — 'Feasible at scale?'
- Competitor — 'How quickly can we copy?'
- Regulator — 'What compliance is needed?'
- 2Each question → grade the answer
- Strong — keep as-is
- Weak — patch/strengthen the related section
- No answer — flag as gap, document for later
- 3Finalize
prfaq-{project}.md→ invokebmad-help→ suggest create-prd
why- ·Stress test = attempting to break the product on paper before breaking it in reality
- ·If the PR/FAQ withstands 10 adversarial questions, the concept has resilience
- ·Gap documentation → research backlog for Phase 2
how- ·AI cycles through 5 personas, 2-3 hard questions per persona
- ·Iterate until all answers are ≥ 'strong' or gaps are explicitly documented
- ·Adversarial review pattern reused in Phase 4 Code Review (04·F)
- 1
bmad-prd skill — Create / Validate / Update. Essential Spine (§0–§9 always) + Adapt-In Menu (7 groups by concern). Adds .memlog.md audit trail and a Reviewer Gate with HTML report. The former create/edit/validate PRD skills are now deprecated forwarding shims.- 01Discovery phaseSource scan → Brain dump → Artifact scan → Stakes calibration (L0–L4, Hobby → Regulated) → Working mode (Fast / Coaching).what
5 discovery stages before drafting begins.
- 1Source scan — auto-scan
planning_artifacts/; surface paths, user confirms before reading - 2Brain dump — user provides all context in one pass; subagent web research spawns in parallel
- 3Artifact scan — parallel subagents extract decision signals; parent receives relevance-filtered digest
- 4Stakes calibration — Hobby/solo · Internal tool · Consumer launch · Regulated → sets depth for every §, classifies L0–L4
- 5Working mode — Fast Path (batch gaps → draft + [ASSUMPTION] tags) or Coaching Path (section by section); switchable at any point
why- ·Stakes calibration sets NFR thresholds, compliance flags, accessibility floor — determines depth of every §
- ·Single brain-dump pass prevents scattered interruptions throughout the session
- ·Early working-mode selection lets PM control pace vs. quality trade-off
how- ·Invoke:
/bmad-agent-pm → CP(Create),VP(Validate),EP(Edit) - ·Stakes probe: 1 question, 4 options — no spam; single calibration determines the whole session depth
- 1
- 02§1 · Vision Always2–3 paragraphs: what the product is, who it serves, why it matters. Blue Ocean lens for launch-level. Pyramid Principle structure.what
2–3 paragraphs compelling enough to stand alone. 3–5 year direction for launch-level; tighter for internal tools. Blue Ocean lens: Eliminate · Reduce · Raise · Create vs. current alternatives.
whyEngineering needs 'why now' and 'why us' — not just 'what'. A decision-maker must understand the go/no-go rationale from §1 alone.
how- ·Ask: 'Where do you want this product to be in 3 years?' — connect to company strategy
- ·Pyramid Principle: lead with conclusion → supporting argument → detail
- ·
- 03§2 · Target User Always§2.1 JTBDs (functional/emotional/social/contextual) · §2.2 Non-Users v1 · §2.3 Key User Journeys (UJ-1…N, named personas, 3–5 beats each).what
- 1§2.1 JTBDs — 4 types: functional · emotional · social · contextual. Format: 'When [situation], I want [motivation] so I can [outcome].'
- 2§2.2 Non-Users (v1) — who this is explicitly NOT for. Omitting this invites scope creep.
- 3§2.3 Key User Journeys — named-persona narratives, numbered UJ-1…N. Each UJ: persona + context → entry → path (3–5 beats) → climax (value delivered) → resolution. FRs reference UJ-IDs inline.
whyNamed personas ('Linh, bookkeeper at XYZ') are mandatory — 'power user' is not actionable. Non-Users establish explicit out-of-scope audience.
how- ·Interview ≥1 real user per persona before the session
- ·Narrate a real session → structure into UJ-N form
- 1
- 04§3 · Glossary Always · New v6.9Every domain noun defined once. No synonyms elsewhere. Agent builds automatically while drafting. Downstream workflows use Glossary terms as anchors.what
One definition per term — if §4 introduces a new domain noun, it gets added to the Glossary in the same editing pass. New in v6.9: downstream workflows extract source content using Glossary terms as anchors.
whyConsistent vocabulary prevents 'is a client the same as an account?' across UX, architecture, and dev teams.
howAgent builds automatically — no dedicated session. Drop when entire readership shares vocabulary. Keep for mixed audiences or regulated domains with precise legal terms.
- 05§4 · Features AlwaysFRs numbered globally FR-1…N. Format: [Actor] can [capability] [conditions]. Realizes UJ-X. Consequences: testable. [ASSUMPTION] inline. Tech choices → addendum.md.what
Features grouped by function; each group opens with a behavioral narrative. FRs nested under feature, numbered globally.
[ASSUMPTION: ...]tags embedded inline when agent inferred without confirmation.why- ·Numbered FRs = backbone of downstream traceability — architecture, stories, QA all reference FR-IDs
- ·Testable Consequences replace acceptance criteria — each FR must pass without ambiguous interpretation
how- ·Ask: 'If we could only ship 5 features, which 5 would you refuse to launch without?'
- ·Write 'what', not 'how' — tech choices belong in
addendum.md
- ·
- 06§5 · Non-Goals AlwaysWhat the product is NOT and will NOT do in v1. Three types: in-scope / out-of-scope / future scope. Sign-off required before drafting FRs.what
Bulleted.
[NON-GOAL for MVP]callouts inline within §4 cover deferred items within features; this section captures the broader 'we are not building X / we are not becoming Y' statements.whyPrevents 'let me add this nearby thing' scope creep at every level. Unwritten out-of-scope items become future misunderstandings.
how- ·Ask: 'What features have we discussed that are definitely NOT happening in this release?'
- ·Do not proceed to drafting FRs without explicit sign-off on scope boundaries
- ·
- 07§6 · MVP Scope AlwaysIn Scope (bulleted) + Out of Scope (each item: reason + defer target: v2, v3, research track). [NOTE FOR PM] for emotionally load-bearing items.what
- ·In Scope — bulleted, crisp
- ·Out of Scope for MVP — each item has a 1-line reason + explicit defer target (v2, v3, research track)
whyMVP ≠ Phase 1: MVP = minimum usable + shippable ('would real users pay for / rely on this?'). A Phase 1 containing 80% of features is not a phase — it's the whole product with a label. False phasing creates false deadline pressure.
how- ·Ask: 'If we had to ship in half the time, what would we cut?'
- ·Dependency blockers and half-time cuts reveal the true MVP
- ·
- 08§7 · Success Metrics Always · Counter-metrics new v6.9Primary (core thesis) + Secondary (supporting) + Counter-metrics (what NOT to optimize — prevents gaming). Each SM cross-references FR(s) it validates. Baseline required.what
3 tiers: Primary (validates core thesis) · Secondary (supporting signals) · Counter-metrics (prevents gaming the primary — new in v6.9, load-bearing). Length scales with stakes.
why- ·Counter-metrics tell the architect what NOT to optimize. Example: maximizing completion rate could tank accuracy.
- ·Primary must be an outcome metric (user value), not an output metric (features shipped)
how- ·Every metric needs a baseline: 'from 45 minutes to under 5 minutes'. 'Reduce time significantly' is not a metric.
- ·Ask: 'What number would you look at 6 months after launch to decide if this was worth building?'
- ·
- 09§8 · Open Questions AlwaysNumbered list of unresolved issues. Become future tickets or follow-up research — not silent gaps in the document.what
Numbered list. Everything still unknown, unresolved in the session. Agent flags every ambiguity here rather than silently assuming.
whyExplicit open questions → tracked tickets. Silent gaps → late-stage surprises.
howCollected during drafting; surfaced before Finalize for PM to triage.
- 10§9 · Assumptions Index AlwaysEvery [ASSUMPTION] tag embedded inline surfaced and indexed here. Every index entry must match an inline tag — mismatch flagged as Critical in Reviewer Gate.what
Index of every
[ASSUMPTION: ...]tag from throughout the PRD. Agent collects during drafting; PM confirms explicitly before Finalize.whyUntested assumptions = Phase 3-4 surprises. The index makes them visible and actionable.
howReviewer Gate cross-checks index vs. inline — mismatch is flagged as a Critical finding.
- 11Adapt-In Menu Optional7 optional groups, identified by Concern Scan: (1) Cross-cutting NFRs · (2) Consumer/branded · (3) Enterprise · (4) Regulated domains (13 industries) · (5) Developer products · (6) Embedded/hardware · (7) Small-scope stories.what
Concern Scan runs automatically as the agent reads Discovery inputs — identifies concerns this product carries and activates the corresponding clusters.
- 1Group 1 — Cross-cutting: Cross-cutting NFRs, Constraints & Guardrails, Why Now
- 2Group 2 — Consumer/branded: Aesthetic & Tone, Information Architecture, Monetization, Platform
- 3Group 3 — Enterprise: Stakeholders & Approvals, Risk & Mitigations, ROI, Operational Requirements, Integration & Dependencies, Rollout, Data Governance, Audit Trail
- 4Group 4 — Regulated domains: Compliance & Regulatory (13-industry taxonomy)
- 5Group 5 — Developer products: API Contracts, Versioning & Deprecation, Performance Budgets, Language/Runtime Targets
- 6Group 6 — Embedded/hardware: Hardware Constraints, Deployment & Update, Environmental & Reliability
- 7Group 7 — Small-scope: Stories (1–2 story max; light PRD + actionable stories in one doc)
whyEssential Spine §0–§9 = mandatory. Adapt-In = optional sections added only when the product actually carries that concern — no bloating PRD with irrelevant sections.
how- ·Concern Scan is automatic — agent flags applicable clusters from Discovery inputs
- ·When a product carries a concern no cluster names → agent invents the section
- 1
- 12Finalize — 8 stepsDecision log audit → Input reconciliation → Reviewer pass (Critical/High first) → Triage open items → Polish (structure before prose) → External handoffs → Close → on_complete hook.what
- 1Decision log audit — walk
.decision-log.md; confirm each entry captured or explicitly set aside - 2Input reconciliation — subagent per source checks gaps, especially qualitative ideas the FR structure silently drops
- 3Reviewer pass — rubric walker + adversarial reviewer in parallel; findings tiered Critical/High first
- 4Triage open items — [ASSUMPTION] tags, [NOTE FOR PM]; phase-blockers resolved first
- 5Polish —
bmad-editorial-review-structurethenbmad-editorial-review-prose; structural before prose - 6External handoffs — Confluence, Notion, Jira...; surface returned URLs/IDs
- 7Close — set
status: finalin frontmatter; log to.decision-log.md - 8on_complete hook — run if configured in
customize.toml
why- ·Decision log prevents re-debate — 6 months later nobody remembers why option B was rejected
- ·Polish structural before prose — do not polish text that will soon be cut
how- ·Output files:
PRD.md+addendum.md+.memlog.md(append-only audit trail) - ·Common next steps:
bmad-ux·bmad-architecture·bmad-create-epics-and-stories
- 1
- 01Load PRD & contextRead PRD, understand users, FRs and NFRs relevant to UX.what
AI loads PRD and extracts UX-relevant FRs + NFRs before starting design.
- 1AI loads
prd.md, extracts personas + FRs with UI surface - 2Identify NFRs that constrain UX
- Performance — p95 <200ms rules out heavy animations
- Accessibility — WCAG 2.1 AA requires focus states, contrast ratios
- Responsive — mobile-first if mobile traffic >50%
- 3Output: FR → UX implication table (e.g.: FR-03 'user uploads file' → drag-drop + progress + error handling)
why- ·UX is not separate from requirements — NFR constraints rule out certain design patterns
- ·Starting with requirements load prevents sketching wireframes that need to be revised
- ·Performance NFR in particular: 200ms threshold rules out long skeleton screens
how- ·Requirements traceability: FR → UX needs mapping table
- ·AI highlights NFRs that constrain design (mark with ⚠️ in notes)
- ·Personas extracted → primary persona drives primary flow
- 1
- 02User journey mappingMap the journey for each persona — touchpoints, emotions, pain points.what
AI maps user journey for each persona: touchpoints, emotions, pain points, opportunities.
- 1Journey map 5 stages
- Awareness — where users hear about product
- Consideration — evaluation, comparison
- Onboarding — signup → first meaningful action
- Regular use — habitual engagement pattern
- Advocacy — retention, referral, renewal
- 2Each stage: touchpoints + emotion (frustrated/curious/delighted) + pain + opportunity
- 3Flag moments of truth — transitions where users most likely to drop off
why- ·Static screens hide flow — journey reveals transitions where UX problems live
- ·Most dropoff happens at transitions, not at the screens themselves
- ·Peak-end rule: design peak delight moment + strong end moment
how- ·Journey mapping methodology with an emotion curve
- ·Transitions tagged with 'danger level' — high = prioritize design attention
- ·One journey per persona — don't merge (different people, different paths)
- 1
- 03Information architectureInformation structure, navigation, sitemap, content hierarchy.what
AI defines Information Architecture: sitemap, navigation hierarchy, content groupings.
- 1Output of this step
- Sitemap — all screens/pages and relationships
- Navigation structure — top-nav / sidebar / breadcrumb decision
- Content taxonomy — categories, labels, search terms
- 2AI uses card sorting to discover user's mental model
- 3Validate: can the user find feature X in ≤3 clicks?
why- ·Bad IA = user gets lost even when the product has enough features
- ·Research: if users can't find it in 3 clicks, they assume the feature doesn't exist
- ·Defining IA upfront = no costly rework after mockups are polished
how- ·Card sorting: open (user groups) or closed (user fits into pre-set)
- ·Tree testing: validate IA with real tasks
- ·Output: sitemap diagram + nav structure decision rationale
- 1
- 04Wireframes (low-fi)Sketch the main screens with layout and basic structure.what
AI sketches low-fidelity wireframes for 5-10 main screens — gray boxes, no color.
- 1Mobile-first: sketch mobile viewport first, desktop after
- 2Gray rectangles + placeholder text — focus on layout + hierarchy, not aesthetics
- 3Per screen annotations
- Primary action — most important thing user does here
- Navigation — entry/exit points
- Content blocks — type, priority, size estimate
why- ·Lo-fi because high-fi too early = team debates colors instead of debating flow
- ·Lo-fi cheap to discard — wrong → fix in 30 minutes, not 3 days
- ·Mobile-first reveals priorities: limited space forces what is actually important to surface
how- ·Crazy 8s technique: sketch 8 variations in 8 minutes before committing
- ·Gray only — no brand colors, no real images, no real text
- ·Annotation mandatory — no annotation = design intent cannot be shared
- 1
- 05Interaction patternsDefine patterns (form, search, filter, modal...) and micro-interactions.what
AI defines interaction patterns — standardize behavior across product.
- 1Core patterns to define
- Forms — validation timing (on-blur vs on-submit), error display, inline vs summary
- Search — autocomplete, filters, results pagination
- Modals — when to use, dismissible via ESC/backdrop, focus trap
- Notifications — toast / inline / system / email trigger conditions
- Loading states — skeleton / spinner / progress bar — when each
- 2Each pattern: trigger condition + behavior + edge cases
why- ·Pattern reuse = consistency — predictability builds trust
- ·Define once → all forms behave same way; user learns once, applies everywhere
- ·Inconsistent patterns = cognitive load on every screen the user must re-learn
how- ·Reference Material Design / Apple HIG / Carbon for industry-standard patterns
- ·Custom patterns only when the product has a strong reason — document rationale
- ·Pattern library output feeds developer component specs in step 06
- 1
- 06Component specsDescribe component types, states, variants — prepare for dev.what
AI lists component specs — each component has states + variants + behavior.
- 1Atom-level components by Atomic Design
- Atoms — Button, Input, Label, Icon, Badge
- Molecules — FormField, SearchBar, Card, MenuItem
- Organisms — Header, DataTable, Modal, Sidebar
- 2Each component spec
- States — default / hover / active / disabled / error / loading
- Variants — primary / secondary / destructive / ghost
- Accessibility — ARIA role, keyboard, focus visible
why- ·Component spec is the contract between UX and Dev
- ·Underspec → dev guess → inconsistency across screens
- ·Atomic Design (Brad Frost): build small, compose large — reusable and maintainable
how- ·Atomic Design methodology: atoms → molecules → organisms → templates → pages
- ·Each component: visual spec + behavior spec + a11y spec + code stub
- ·Export to Figma component library if the team uses Figma
- 1
- 07Responsive designBreakpoints, mobile/tablet/desktop behaviors, adaptive patterns.what
AI defines responsive behavior: breakpoints + per-breakpoint layout decisions.
- 1Standard breakpoints
- Mobile — <640px: single column, bottom nav
- Tablet — 640-1024px: 2 column, side nav optional
- Desktop — >1024px: multi-column, expanded nav
- 2Per-breakpoint decisions
- Navigation pattern changes (bottom nav → side nav → top nav)
- Content priority (mobile hides secondary actions)
- Component reflow (card stack vs grid)
why- ·70% web traffic mobile — mobile-first is not an afterthought
- ·Mobile-first vs desktop-first determines downstream code structure
- ·Explicit breakpoints prevent developers from picking random values (768? 800? 1024?)
how- ·Progressive enhancement: start mobile, add capabilities for larger screens
- ·Content priority per breakpoint: what shows, what hides, what reflows
- ·Test on real devices — DevTools doesn't capture touch behavior
- 1
- 08AccessibilityWCAG 2.1 AA: keyboard nav, screen reader, contrast, focus states.what
AI specifies accessibility per WCAG 2.1 AA — not a nice-to-have.
- 1Core WCAG 2.1 AA requirements
- Contrast — 4.5:1 for normal text, 3:1 for large text
- Keyboard nav — all interactive elements reachable via Tab
- Screen reader — meaningful labels, alt text, ARIA where needed
- Focus visible — clear focus indicator on all elements
- Error identification — errors described in text, not only color
- Motion — prefers-reduced-motion respected
why- ·Legal requirement in many markets: ADA (US), EAA (EU 2025)
- ·Cost of retrofit 10x cost of build-in
- ·Accessibility improvements help all users — SEO, keyboard power users, mobile
how- ·WCAG 2.1 AA checklist + ARIA patterns
- ·Tools: aXe DevTools, Lighthouse a11y audit
- ·Manual test: keyboard-only navigation + screen reader (NVDA/VoiceOver)
- 1
- 09Design system referencesLink to the design system / tokens, or recommend creating one.what
AI references or proposes building a design system with semantic tokens.
- 1Design tokens — W3C standard
- Color — semantic naming:
color.bg.primarynotcolor.blue.500 - Typography — scale: xs/sm/md/lg/xl, line-heights, weights
- Spacing — 4px base grid, named scale: 1/2/3/4/6/8/12/16...
- Shadow, radius, motion — consistent values
- Color — semantic naming:
- 21 token change propagates everywhere — brand refresh = update 20 tokens, not 200 screens
why- ·Without design system: every screen reinvents wheels
- ·Token-based approach lets 1 change propagate everywhere
- ·Critical for Level 3-4 projects with multiple feature teams
how- ·W3C Design Tokens standard format
- ·Reference: Tailwind, Material, Carbon, Polaris — pick closest to team's stack
- ·Semantic naming: describes role, not appearance
- 1
- 10Finalize spinesSave DESIGN.md + EXPERIENCE.md, hand off to Architecture.what
AI saves the two peer spines —
DESIGN.md(visual identity + tokens) andEXPERIENCE.md(IA, behavior, states, a11y, journeys; references DESIGN.md tokens) — and clearly hands off to Architecture.- 1Spine structure
- User journeys — EXPERIENCE.md, named-protagonist
- Information Architecture — EXPERIENCE.md, sitemap + nav
- Interaction primitives + state patterns — EXPERIENCE.md
- Component patterns — behavioral in EXPERIENCE.md, visual specs in DESIGN.md
- Accessibility floor — EXPERIENCE.md (visual contrast in DESIGN.md)
- Design tokens + brand/style — DESIGN.md
- 2AI invokes
bmad-help→ suggest Architecture (Phase 3) next
why- ·UX spec is an input to Architecture — component list informs frontend framework choice
- ·Accessibility spec informs library choice (a11y-first vs not)
- ·Interaction patterns inform state management approach
how- ·Handoff format standard BMAD — frontmatter complete with stepsCompleted
- ·Wireframes: embed as ASCII/text description if Figma is not available
- ·
bmad-helparg='ux-design' → recommend next: architecture
- 1
bmad-architecture skill (create / update / validate) — lean-spine model producing ARCHITECTURE-SPINE.md as the primary source of truth, with a breadth-coverage rubric and a Reviewer Gate. Supersedes the deprecated fixed 11-step architecture flow.- 01Detect intent & activation modeCreate / Update / Validate — resolved from the conversation and input, not quizzed. Headless runs follow references/headless.md end to end; forwarded calls from the deprecated bmad-create-architecture shim honor pre-resolved fields verbatim.what
Resolves customization, loads
bmm/config.yaml, then detects one of three intents: Create (default), Update an existing spine, or Validate one without changing it.whyA single skill replacing three fixed flows must route correctly on the first turn — misrouting into Create when the user meant Validate wastes the whole session.
howIf a run folder already exists under
{workflow.spine_output_path}, offer to resume from its memlog instead of restarting. If the real ask is requirements/UX/a capability contract/epic breakdown, redirect tobmad-prd,bmad-ux,bmad-spec, orbmad-create-epics-and-storiesinstead. - 02Choose Coaching vs Fast pathCoaching (default) — open-ended elicitation, load-bearing calls shown not silently made. Fast — draft the whole spine with [ASSUMPTION] tags for review.what
Offered as an activation step, in the user's language, before any drafting. Coaching path pulls decisions out of the user with open-ended questions; Fast path infers and tags.
whyElicitation is the value the skill is coaching toward — silently drafting the whole spine defeats the purpose unless the user explicitly wants speed.
howAlso asks, mandatory on both paths: is the spine the only deliverable, or does the user need a purpose-scoped human-facing artifact (team walkthrough, board vision doc) later at Finalize?
- 03Read the input to know the jobSpec package, raw idea, sprawling doc to distill, existing codebase, one feature's slice, or an existing spine to extend/pressure-test — the input's shape decides the job.what
A
SPEC.md+ memlog is the richest, preferred start. Brownfield work investigates real code andproject-context.mdto ratify existing conventions rather than invent new ones.whyThe spine's altitude must mirror what it augments (initiative→features, feature→epics, epic→stories) and stay coherent with whatever level sits below it.
howInheriting a parent spine (e.g. one epic of an existing feature spine): its ADs and paradigm load as binding, read-only Inherited Invariants — only the parent's Deferred items are this run's job.
- 04Establish paradigm & seedLead with a named design paradigm; recommend a verified current starter for greenfield; keep seed (stack, tree, data shape) minimal — only invariants are fixed.what
One test decides what belongs in the spine: could two units built independently choose incompatibly, and is the call non-obvious and a real trade-off? If not, it's Deferred.
whyEverything structural (the seed) is true at cold-start and owned by the code once it exists — over-fixing seed content invites drift between spine and reality.
howVerify any named technology's current version and fit on the web before binding it into the spine.
- 05Log decisions to the memlogEvery decision, constraint, version, assumption, and open question lands as one append-only line in .memlog.md — the run's working memory, not the rendered spine.what
Each surviving decision becomes an
AD-n(stable ID, Binds / Prevents / Rule); a decision that lives only in a diagram is still logged.whyDistilling the spine from a living log — instead of hand-editing the artifact — lets Update runs amend a Rule in place and add the next
AD-nwithout ever renumbering or reusing a retired ID.howWrites go through the shared
memlog.py init/append --type <decision|constraint|version|assumption|question|direction|event>script — never a hand-patch to the spine file. - 06Distill the spineWrite ARCHITECTURE-SPINE.md from the memlog — invariants first, seed minimal, every AD carrying Binds/Prevents/Rule, Deferred naming what it won't decide.what
Sweeps the breadth the altitude owns — every structural dimension (including the operational/environmental envelope: deployment, infra, operations) is decided, deferred, or an open question. No placeholders; nothing invented to fill a gap.
whyA whole dimension left silent is the failure mode this rubric exists to catch — not a clean spine.
howA subagent per load-bearing input reconciles the draft against its source and flags anything the AD structure quietly dropped, before the Reviewer Gate.
- 07Reviewer GateDeterministic lint_spine.py pass + a good-spine rubric walker + every finalize_reviewers lens dispatched as parallel subagents against ARCHITECTURE-SPINE.md.what
At Finalize/Update, clear fixes from the gate are applied directly. Under the standalone Validate intent, the same gate instead produces a bespoke HTML report and hands findings back to the user.
whyScaled to stakes — a small feature slice gets a light pass; a platform-wide spine gets the full lens set.
howOpen questions and [ASSUMPTION] tags triage into blockers (resolved one at a time before handoff) and deferred items (logged with a revisit condition).
- 08Renderings, close & handoffOptional human-facing artifact scoped to the up-front purpose; set status: final; recommend bmad-spec, then bmad-create-epics-and-stories or bmad-create-story.what
The spine is the build deliverable. Optional extras — an interactive HTML+SVG walkthrough deck, a fuller solution design doc, a C4 set, a team/epic split view — are built only if the up-front purpose called for one.
whyADIDs stay stable across handoff so downstream skills (epics, stories, spec) can cite them without drift.howSets the spine's own frontmatter
status: final, logs amemlog.py append --type event --text "spine finalized", then leads withbmad-spec(adopt/refresh the spine as a spec companion) beforebmad-create-epics-and-storiesorbmad-create-storyat epic altitude.
- 01Load planning docsRead PRD, architecture, epics, UX spines (DESIGN.md + EXPERIENCE.md) — gather all planning artifacts.what
PO agent loads all planning artifacts and builds inventory.
- 1Load documents
prd.md+addendum.md+decision-log.mdARCHITECTURE-SPINE.md+ADRs/*.mdepic-{N}.mdall epicsDESIGN.md+EXPERIENCE.mdif available
- 2PO builds artifact inventory: file list + version + last-updated
- 3Flag missing artifacts before starting review
why- ·Readiness check is the gatekeeper before Phase 4
- ·PO ≠ PM — independent reviewer reduces self-confirmation bias
- ·BMAD tip: use different LLM for validation — same LLM finds its own writing consistent
how- ·PO profile: skeptical, thorough, user-advocate perspective
- ·Artifact inventory: date comparison (architecture newer than PRD = check for drift)
- ·Missing artifacts = automatic Concerns flag
- ·Version pin: record the commit/hash of each doc to ensure the review uses the correct version
- 1
- 02PRD quality reviewCheck completeness, clarity, FRs/NFRs complete and non-contradictory.what
AI reviews PRD: completeness + consistency + FR quality.
- 1Completeness check
- Every required section present and non-empty?
- Every FR has ID + title + description + priority + user story?
- Every NFR has a metric + threshold + measurement method?
- Exec summary present and covering all major points?
- 2Consistency check
- Exec summary match body detail?
- FR-IDs referenced consistently (FR-03 in one place, FR-3 elsewhere = error)?
- NFRs do not contradict FRs?
why- ·PRD bug most expensive when fixed early — propagates downstream to architecture and stories
- ·1 missing FR = 1 missing feature = customer complaint
- ·1 vague NFR = battle between dev and stakeholder post-ship
how- ·Template-driven check: section X exists? Content X has substance?
- ·ID consistency regex: normalize all references before check
- ·Output: checklist with PASS/FAIL/CONCERN per item
- 1
- 03Architecture reviewADRs have clear rationale, NFRs are addressed, risks have mitigations.what
AI reviews Architecture: ADR quality + NFR coverage + risk register.
- 1ADR quality
- Every ADR has complete Context/Decision/Consequences/Alternatives?
- Status field set (Proposed/Accepted)?
- Do all tech stack decisions have ADRs?
- 2NFR coverage
- Each NFR from PRD has a component/pattern addressing it in the architecture?
- Does the 99.9% uptime NFR have an HA topology?
- Does the security NFR have a threat model?
- 3Risk register present? Do the top 5 risks have mitigations?
why- ·Architecture review is often skipped — 'trust the architect'
- ·Independent review catches: missing ADR (decision not documented), under-addressed NFR
- ·NFR-architecture gap = failure surprise in production
how- ·ATAM-light: each NFR → find addressing component → rate adequacy
- ·Tech stack alignment against team skills check
- ·Vendor lock-in risk graded
- 1
- 04Cohesion checkPRD ↔ Architecture: every FR has a solution, every decision is grounded in PRD.what
AI run cohesion check: PRD ↔ Architecture traceability.
- 1Build traceability matrix
- Rows = FRs from PRD
- Cols = architecture components
- X = component addresses FR
- 22 types of issues
- Empty row — FR has no component → architecture incomplete
- Empty column — component not justified by any FR → over-engineering
- 3Flag all gaps as Blocker if FR = Must priority
why- ·Cohesion gap is a silent killer — PRD says 'users upload files', architecture has no FileService
- ·Discovered after shipping = emergency fix, downtime risk
- ·Over-engineering = wasted sprint capacity + maintenance burden
how- ·Same traceability matrix technique as 03·B step 01
- ·PO perspective: 'which feature will users miss?' → check corresponding FR row
- ·ADR cross-check: every architecture decision traces back to a PRD requirement?
- 1
- 05Epic quality reviewValidate against best practices: user value, independence, dependencies.what
AI reviews epics: INVEST compliance + AC quality + dependency graph.
- 1Per story INVEST check
- Independent? (no hard block on adjacent story?)
- Estimable? (dev can estimate without major uncertainty?)
- Small? (≤3 days dev work?)
- Testable? (≥2 AC defined?)
- 2AC quality check
- Stories with 0-1 AC → insufficient spec flag
- Stories with >7 AC → too large, split flag
- Gherkin format: Given/When/Then present?
- 3Dependency graph cycle check: A→B→A = circular dependency = blocker
why- ·Bad stories at this gate = bad sprint — sprint failure more expensive
- ·Story without clear AC → code no one can verify
- ·Circular dependency → sprint never completes
how- ·Automated INVEST scoring: count AC, check size estimate present
- ·Cycle detection: graph traversal (DFS cycle check)
- ·Output: per-story scorecard with specific issues
- 1
- 06Final assessmentConclusion: PASS, CONCERNS (minor fixes needed), or FAIL (must go back).what
AI concludes with a binary verdict and saves the readiness report.
- 1Verdict criteria
- PASS — 0 Blockers + ≤3 Concerns → ready for Phase 4
- NEEDS REVIEW — 0 Blockers + >3 Concerns → proceed with documented caveats
- FAIL — ≥1 Blocker → must rework before Phase 4
- 2Report structure
- Scorecard — per-section PASS/CONCERNS/FAIL
- Blockers — must-fix list with owner
- Concerns — should-fix list with rationale
- Verdict — go/no-go with reasoning
why- ·Binary gate (go/no-go) is the critical output
- ·Without explicit gate, project drifts into Phase 4 with planning gaps → sprint chaos
- ·CONCERNS state can still proceed, but with documented risks — informed decision
how- ·Blocker definition: issue that WILL cause Phase 4 failure if not fixed
- ·Concern definition: issue that MIGHT cause problems but manageable
- ·Save:
readiness-report.mdwith timestamp and verdict
- 1
- 01Detect mode & prerequisitesDetermine System-Level vs Epic-Level mode; confirm PRD, ADR, and architecture doc are available; halt with a clear message if required inputs are missing.
- 02Load context & knowledge baseExtract tech stack, integration points, and NFRs from inputs; load tiered knowledge fragments just-in-time (core always, extended on-demand) for 40–50% context saving vs full load.
- 03Testability reviewEvaluate architecture for controllability (can we set preconditions?), observability (can we verify outcomes?), and reliability (are test results deterministic?); classify ASRs as Actionable or FYI.
- 04Risk assessment & prioritizationScore 6 risk categories (TECH/SEC/PERF/DATA/BUS/OPS) with P×I matrix (1–9 scale); assign P0–P3 priority to each test scenario; high-risk items (score ≥ 6) require mandatory mitigation with owner and timeline.
- 05Generate output documentsWrite test-design-architecture.md (testability concerns + risk mitigations for Dev/Arch teams), test-design-qa.md (coverage plan + execution recipe for QA team), and BMAD handoff document bridging to epic decomposition.
- 01Preflight checksRead config, auto-detect stack from project manifests (package.json, pyproject.toml, pom.xml, go.mod); verify no conflicting framework exists; confirm write permissions; save detected context to progress file.
- 02Framework selectionFrontend: default to Playwright, switch to Cypress only if criteria met. Backend: Python→pytest, Java/Kotlin→JUnit 5, Go→go test, C#→xUnit, Ruby→RSpec, Rust→cargo test. Announce selection with rationale.
- 03Scaffold frameworkGenerate complete directory tree, framework config, .env.example, fixtures with mergeTests pattern (auto-cleanup hooks built in), Faker-based data factories, sample tests, and helpers. Adaptive execution: parallel agent-team or sequential based on runtime capability.
- 04Documentation & scriptsCreate tests/README.md covering setup, running, debugging, architecture, best practices, and CI notes; add idiomatic test commands to package.json, Makefile, or pyproject.toml (minimum: "test:e2e" for frontend).
- 05Validate & summarizeCheck all checklist items (directory structure, config correctness, fixtures/factories, docs/scripts); fix any gaps before marking complete; report framework selected, all artifacts created, and recommended next steps.
- 01Load PRD + architectureRead both to understand requirements and existing technical decisions.what
AI loads PRD + ARCHITECTURE-SPINE.md and builds FR ↔ component traceability map.
- 1AI loads both documents in parallel
prd.md— FRs, NFRs, personas, success metricsARCHITECTURE-SPINE.md— components, data model, API design, patterns
- 2Build traceability matrix
- Rows = FRs from PRD
- Cols = architecture components
- X = 'this component addresses this FR'
- 3Flag gaps: FR has no component (architecture incomplete) + component has no FR (over-engineering)
why- ·v6 improvement: stories are created AFTER architecture — stories now reference specific patterns
- ·Stories created before architecture miss tech context → rework mid-sprint
- ·Traceability surfaces gaps BEFORE sprint capacity is committed
how- ·Traceability matrix: rows = FRs, cols = components, X = mapping
- ·Empty row = FR uncovered → must address in architecture
- ·Empty col = component not justified → over-engineering, remove or justify
- 1
- 02Group FRs into epicsGroup related FRs into epics — each epic carries independent value.what
AI groups FRs into epics — each epic is a user-deliverable capability.
- 1AI clusters FRs by user capability, not by technical layer
- 2Validate per epic
- User-deliverable: can it be demonstrated standalone?
- Independent value: do users benefit if only this epic ships?
- Reasonable scope: 2-6 weeks development?
- 3Bad vs good epic naming
- Bad: 'Backend', 'Frontend' — technical layers are not capabilities
- Good: 'User onboarding', 'Payment processing', 'Admin dashboard'
why- ·Tech-layer epics do not deliver value individually
- ·User-centric epics deliver value per epic ship — stakeholders see progress
- ·Validate 'demo standalone' as the acid test for epic independence
how- ·Capability mapping: each FR maps to exactly 1 epic
- ·Epic size check: one epic too large (>6 weeks) → split; too small (<1 week) → merge
- ·Validate against the product vision: each epic must tell a capability story
- 1
- 03Epic orderingDetermine epic order by dependencies, risk, and value.what
AI orders epics by WSJF (Weighted Shortest Job First).
- 1WSJF scoring per epic
- User-Business Value — direct revenue/UX impact (1-10)
- Time Criticality — cost of delaying (1-10)
- Risk Reduction — reduces unknowns if done early (1-10)
- Job Size — estimated weeks (1-10, higher = larger)
- WSJF = (Value + Criticality + Risk) / Size
- 2High-risk, high-value epic → ship early (fail fast while resources remain)
- 3First epic must be independent (no dependencies) — unblock parallel work
why- ·Shipping in the wrong order = waste — first 4 sprints does not deliver demoable value = stakeholder frustration
- ·Risky epics early: if they fail → pivot while still have runway
- ·Value-first: demonstrate value early → stakeholder confidence + feedback
how- ·SAFe WSJF framework
- ·Score subjectively if exact data is unavailable — relative ranking is enough
- ·Visualize: table with scores + bar chart so stakeholders can see the rationale
- 1
- 04Story breakdownSplit epics into stories using INVEST principles.what
AI splits each epic into INVEST stories.
- 1INVEST checklist per story
- Independent — do not block each other when possible
- Negotiable — scope can be adjusted
- Valuable — clear user/business benefit
- Estimable — effort can be estimated
- Small — 1-3 days of dev work is ideal
- Testable — AC can be verified
- 2SPIDR splitting patterns when the story is too large
- Spike — separate exploration/research story
- Path — split by user journey paths
- Interface — split by UI vs API vs integration
- Data — split by data variations
- Rules — split by business rules
why- ·Story too large → uncertainty → wrong estimates → sprint slips
- ·Story too small → overhead (PR review, deploy) > value
- ·INVEST sweet spot: 1-3 days reduces risk and enables fast feedback
how- ·INVEST checklist per story — fail any = must split/reshape
- ·SPIDR patterns: most used = Interface split (API vs UI separate stories)
- ·Story size: count by AC — story with >7 AC is usually too large, split
- 1
- 05Acceptance criteriaWrite AC in Given/When/Then format for each story, with enough detail to test.what
AI writes Acceptance Criteria in BDD/Gherkin format — executable specification.
- 1Gherkin format
- Given — precondition (state before the action)
- When — action user takes
- Then — expected outcome (verifiable)
- 2Coverage per story
- 1+ AC for happy path
- 1-2 edge cases (boundary values, empty states)
- 1 error path (what happens when it fails)
- 33-7 AC per story — under 3 = insufficient spec; over 7 = story too large
why- ·Vague AC ('works correctly') → endless debate about done
- ·BDD format forces precondition + trigger + outcome — do not miss edge cases
- ·Executable AC: if using Cucumber/Playwright, tests write themselves from AC
how- ·Each AC independently testable — do not chain multiple outcomes into one AC
- ·Edge cases from INVEST size estimate: 'what are the boundaries of this story?'
- ·Error paths mandatory: 'What does the user see when it fails?'
- 1
- 06Story dependenciesIdentify dependencies between stories and mark blockers.what
AI builds a dependency graph between stories.
- 13 dependency types
- Hard blocker — story B cannot start until A is done
- Soft dependency — B is easier if A is done, but B can still start
- Independent — no dependency
- 2AI identifies the critical path — longest blocker chain
- 3Flag stories with >3 dependencies — signal must split or re-sequence
why- ·Dependency graph drives sprint planning and parallel work
- ·Critical path = minimum sprint duration — visualize it for the SM
- ·Unresolved dependencies = sprint blocker surprise
how- ·Build graph: nodes = stories, directed edges = blocker
- ·Critical path analysis: longest path = minimum sprint duration
- ·Parallel opportunities: stories with no outgoing dependencies = safe to parallelize
- ·Cross-epic dependency = red flag: story depending on another epic breaks sprint boundaries
- 1
- 07Finalize epic filesSave epic-{N}.md for each epic with all stories inside.what
AI save
epic-{N}.mdper epic — stories nested inside.- 1File structure
- 1 file per epic:
epic-01-user-onboarding.md - Frontmatter: epic ID, status, owner, WSJF score
- Stories inside: S1, S2, S3... with full AC and technical context
- 1 file per epic:
- 2Sharding for large epics
- Option:
stories/epic-01/story-01.mdseparated - Use when epic has >10 stories — avoid one huge file
- Option:
- 3AI invokes
bmad-help→ suggest Readiness Check (03·D) next
why- ·1 file 1 epic = just-in-time loading — Phase 4 loads only 1 epic at a time
- ·Full story content in the epic file = dev agent reads only the epic, no PRD/architecture required
- ·Separation ensures minimal context window per sprint
how- ·Frontmatter:
epicId, status, owner, sequence, wsjfScore - ·Stories numbered S1, S2 within epic — globally unique via
epic-01-S1 - ·
bmad-helprecommend: Readiness Check (03·D) before Phase 4
- 1
- 01Preflight checksVerify Git repo exists, detect stack, confirm test framework config is present, run local tests (halt with "Fix failing tests before CI setup" if they fail); detect CI platform from existing config files or git remote URL; default to GitHub Actions.
- 02Generate CI pipelineSelect platform-specific template; configure all stages (lint → test → contract-test → burn-in → report); set up parallel sharding (4 shards via matrix strategy); route all user-controlled contexts through env: intermediaries to prevent script injection.
- 03Configure quality gates & notificationsEnable burn-in loop (10× iterations, frontend/fullstack by default; backend skipped as deterministic); set quality thresholds (P0=100%, P1≥95%); configure can-i-deploy contract gate if Pact is enabled; wire Slack/email failure notifications with artifact links.
- 04Validate & summarizeValidate generated config against the full checklist (config path, stages, sharding, burn-in, artifacts, secrets); fix any gaps before marking complete; write completion report with config path, enabled stages, required secrets, and next steps (set secrets, push, run pipeline).
- 01Parse epic filesScan epic files and extract epic numbers plus story IDs/titles from headers in the body.what
Scan epic files and extract epic numbers plus story IDs/titles from headers in the body.
why· Sprint planning needs ALL epics and stories to build complete status tracking
how· Whole doc takes priority over sharded docs when both exist
- 02Build status structureCreate a flat development_status map with default statuses for epics, stories, and retrospectives.what
Create a flat development_status map with default statuses for epics, stories, and retrospectives.
whyCreate a flat development_status map with default statuses for epics, stories, and retrospectives.
how· Order: epic → that epic's stories → retrospective → next epic
- 03Apply status detectionCheck real story files, upgrade status when files exist, and never downgrade.what
Check real story files, upgrade status when files exist, and never downgrade.
whyCheck real story files, upgrade status when files exist, and never downgrade.
how· Check path: {story_location}/{story-key}.md
- 04Generate sprint-status.yamlWrite the flat development_status map with metadata header to the file.what
Write the flat development_status map with metadata header to the file.
whyWrite the flat development_status map with metadata header to the file.
how· Output path: {implementation_artifacts}/sprint-status.yaml
- 05Validate + ReportCheck completeness and report totals before finishing.what
Check completeness and report totals before finishing.
why· Artifacts drive state — YAML persists across sessions, every agent reads the same source
howValidate completeness and report totals. Validate that every epic/story/retro has an entry, there are no extra entries, status values are valid, and YAML syntax is correct
- 01Determine target storyUse a provided story path or auto-discover the first backlog story from sprint-status.yaml.what
Read the full sprint-status.yaml, select the first story entry with status
backlog, or parse the user-provided story id/path.whyCreate Story is just-in-time per story. Preparing stories in bulk makes later stories stale after earlier implementation learnings.
howIf this is the first story in an epic, move the epic from
backlogtoin-progress. Halt if no eligible backlog story exists. - 02Load and analyze core artifactsLoad epics, PRD, architecture, UX, previous story learnings, and recent git patterns selectively.what
Extract the target story's user statement, ACs, dependencies, constraints, and any context from prior story files or recent commits.
whyThe story file must contain everything downstream agents need without rereading full planning docs.
howLoad epics in full; load PRD, architecture, and UX sections selectively for the chosen story.
- 03Extract architecture guardrailsCapture tech stack, file locations, naming conventions, API contracts, data schemas, security, and regression constraints.what
Read every UPDATE file the story is expected to touch and document current behavior, intended change, and what must be preserved.
whyWrong file locations, stale APIs, and missed existing behavior are the common failure modes Create Story prevents.
howMake the guardrails explicit in Dev Notes so BMAD Dev Story inherits them directly during implementation.
- 04Research latest technical specificsCheck current library versions, changelogs, security notes, deprecations, and best practices relevant to the story.what
Research only the APIs, frameworks, and external services the selected story will interact with.
whyLLM memory may be stale; story creation is the last cheap point to catch breaking changes or vulnerable patterns.
howEmbed specific versions, API notes, migration cautions, and security patches in the story file.
- 05Create comprehensive story fileWrite the complete story file from template and set Status = ready-for-dev.what
Synthesize requirements, ACs, dev notes, architecture compliance, test expectations, and current-tech notes into
{story_key}.md.whyThe output becomes the canonical story input for Dev Story and every downstream implementation gate.
howInitialize from template, fill every required section, and save clarification questions for the end instead of interrupting the workflow.
- 06Update sprint status and finalizeValidate the story file, update sprint-status.yaml to ready-for-dev, preserve comments, and report the handoff.what
Run checklist validation and update only the matching story status plus
last_updatedin sprint-status.yaml.whyready-for-devmeans the story is prepared for Dev Story but implementation has not started.howReport story id, path, status, and next step: Dev Story implementation.
- 01Load story and epic contextRead the selected story, parent epic, architecture, project context, and approved spec.what
Read the selected story, parent epic, architecture, project context, and approved spec.
whyCoverage must be anchored to actual requirements and architecture constraints.
howInventory functional requirements, scenarios, NFRs, and integration boundaries.
- 02Assess riskScore probability and impact for each scenario or requirement group.what
Score probability and impact for each scenario or requirement group.
whyThe highest-risk behavior deserves the strongest evidence.
howMap risks into P0-P3 priorities and record the rationale.
- 03Choose test levelsAssign unit, integration, contract, E2E, and NFR checks to the right risks.what
Assign unit, integration, contract, E2E, and NFR checks to the right risks.
whyA balanced pyramid gives confidence without creating a brittle suite.
howUse unit tests for logic, integration for boundaries, E2E for critical journeys, and NFR checks for quality attributes.
- 04Generate coverage planWrite the per-story coverage strategy and evidence expectations.what
Write the per-story coverage strategy and evidence expectations.
whyThe plan becomes the test architecture contract for ATDD, automation, and release traceability.
howDocument test levels, fixtures, data, priority, and owner.
- 05Validate and hand offCheck for uncovered or untestable requirements and hand gaps to ATDD, Dev, or PM.what
Check for uncovered or untestable requirements and hand gaps to ATDD, Dev, or PM.
whyGaps found here are cheaper than gaps found at release gate.
howFlag missing ACs, ambiguous requirements, and coverage decisions needing human approval.
- 01Read acceptance criteriaExtract behaviors, examples, error paths, and observable outcomes.what
Extract behaviors, examples, error paths, and observable outcomes.
whyATDD must test user-visible behavior, not implementation guesses.
howMap each AC to one or more executable examples.
- 02Write behavior scenariosTranslate acceptance criteria into Given/When/Then or equivalent spec cases.what
Translate acceptance criteria into Given/When/Then or equivalent spec cases.
whyScenario language exposes missing preconditions and expected outcomes.
howName scenarios after behavior, not functions.
- 03Generate skipped scaffoldCreate skipped/disabled acceptance tests, fixtures, and helper stubs without implementing production behavior.what
Create skipped/disabled acceptance tests, fixtures, and helper stubs without implementing production behavior.
whyThe scaffold defines the expected red signal without letting implementation start before ATDD handoff is complete.
howKeep assertions meaningful, mark each scaffold as intentionally skipped/disabled, and record the command that should turn red after activation.
- 04Prepare red checklistRecord which scaffolds Dev must activate, the expected red failure, and the command to run.what
Record which scaffolds Dev must activate, the expected red failure, and the command to run.
whyATDD owns acceptance intent and scaffold quality; Dev owns the red observation after activation inside BMAD Dev Story's TDD loop.
howDo not require Murat to turn tests red here. The checklist tells Dev Story exactly what to activate and what red signal to expect.
- 05Hand off to Dev StoryAttach scaffold paths, activation checklist, test commands, and expected red failures to the story file for Dev Story.what
Attach scaffold paths, activation checklist, test commands, and expected red failures to the story file for Dev Story.
whyDev Story must consume ATDD output before implementation starts so coding starts from an acceptance contract.
howLink tests to tasks and acceptance criteria, then pass the checklist directly into the story file Dev Story reads.
- 01Find & load storyAuto-discover the story with status=ready-for-dev in sprint-status.yaml, or use the story path the user provides; read the FULL story file.what
Determine which story to implement and read its entire content.
- ·User provides a story path → use it directly
- ·No path given → read the FULL sprint-status.yaml → find the first N-N-name entry with status
ready-for-dev - ·Read the COMPLETE story file: AC, Tasks/Subtasks, Dev Notes, Dev Agent Record, File List, Change Log, Status
- ·Load
project-context.mdif it exists - ·HALT if no story is found or the story file is not accessible
why- ·The story file is the single source of truth — every implementation decision comes from it
- ·Dev Notes carry architecture guardrails and learnings from Create Story — skipping them loses critical context
how- ·Read the FULL sprint-status.yaml (every line) → filter N-N-name entries → take the first with status=ready-for-dev
- ·If none is ready-for-dev → ask the user to run Create Story, validate it, or provide a story path
- ·
- 02Detect review continuationCheck whether the story already went through Code Review — if so, extract pending follow-up items to process first.what
- ·Scan the story file: does it have a "Senior Developer Review (AI)" section?
- ·If yes →
review_continuation = true: extract the review outcome, count unchecked [AI-Review] items, save the list of pending items - ·If no →
review_continuation = false: fresh start
why- ·Review follow-ups must be processed before regular tasks — the reviewer already identified specific issues
- ·Skipping detection risks silently skipping high-severity review findings
how- ·Scan the "Review Follow-ups (AI)" subsection under Tasks/Subtasks
- ·Count High / Med / Low severity items still unchecked
- ·
- 03Mark in-progressCapture baseline_commit into the story's YAML frontmatter; update sprint-status.yaml: ready-for-dev → in-progress.what
- ·Only run if the current status =
ready-for-devand the story YAML has nobaseline_commityet - ·Run
git rev-parse HEAD→ save into the story's YAML frontmatter asbaseline_commit: <sha> - ·Update sprint-status.yaml: story_key →
in-progress, updatelast_updated
why- ·
baseline_committells Code Review which commit the diff should start from - ·Status
in-progresssignals to the team that the story is actively being implemented
how- ·If git is not available →
baseline_commit: NO_VCS - ·If the YAML frontmatter already has
baseline_commit→ do not overwrite it - ·Update sprint-status.yaml: rewrite the FULL file, preserving ALL comments and structure
- ·
- 04Implement task — RED·GREEN·REFACTOR loopRepeat per task: write failing tests (RED) → implement minimal code (GREEN) → refactor → validate → mark [x] → next task, without stopping midway.what
Loop through each task in the order listed in the story's Tasks/Subtasks:
- ·RED: write failing tests → confirm tests fail before writing code
- ·GREEN: implement the minimal code needed to pass — nothing outside task scope
- ·REFACTOR: improve code structure; tests stay green after refactor
- ·Validate: run the full test suite, check the AC tied to the task
- ·Mark the task
[x]only once all its tests actually exist and pass 100% - ·If review_continuation = true: process the [AI-Review] items before regular tasks
why- ·TDD forces the behavior to be articulated before an implementation is chosen — avoids over-engineering
- ·NEVER stop for "milestones," "significant progress," or "session boundaries" — only HALT on a specific blocking event
- ·NEVER mark [x] unless tests actually pass — no lying or cheating about test status
how- ·HALT if: a dependency outside the story spec is needed, 3 consecutive implementation failures occur, or required config is missing
- ·Update the File List after each task: ALL new/modified/deleted files (relative paths)
- ·Record completion notes in the Dev Agent Record after each task completes
- ·
- 05Story completion & DoD validationVerify all tasks are [x], run full regression, validate the Definition of Done, update status: in-progress → review.what
Gate before the story enters the review queue:
- ·Re-scan the entire story document: verify ALL tasks and subtasks are already
[x] - ·Run the full regression suite — don't skip it
- ·DoD checklist: all AC satisfied, unit/integration/e2e tests added, linting passes, File List complete, Dev Agent Record has notes, Change Log has a summary
- ·Update story Status →
review - ·Update sprint-status.yaml: story_key →
review, preserving ALL comments
why- ·DoD validation before signaling review avoids wasting the reviewer's time on incomplete stories
- ·Status
review(notdone) — the story waits for Code Review approval before it's done
how- ·HALT if: any task is not yet [x], regression fails, File List is missing, or DoD fails
- ·YAML update: change only status and last_updated — don't rewrite the whole file
- ·
- 06Communication & handoffReport completion tailored to user_skill_level; suggest running Code Review with a different LLM.what
- ·Summary: story ID, story key, title, key changes, tests added, files modified
- ·Provide the story file path and current status (review)
- ·Ask whether the user needs anything explained (per user_skill_level)
- ·Suggest next steps: run Code Review, verify ACs, check sprint status
why- ·The user needs to know the story is done and may have questions about implementation decisions
- ·Code Review should use a different LLM to avoid blind spots from the model that already implemented the story
how- ·Tailor the explanation to
user_skill_level— junior needs more detail, senior needs trade-offs - ·Tip: "For best results, run Code Review using a different LLM than the one that implemented this story"
- ·
- 01Gather contextFind the review target via a 5-tier cascade; construct the diff; determine review_mode: full (spec/story file present) or no-spec.what
Find the review target via a 5-tier cascade (stop at the first tier that matches):
- ·Tier 1: explicit arg — PR, commit SHA, branch, spec file, diff-mode keyword (staged/uncommitted/branch diff/commit range)
- ·Tier 2: recent conversation context — scan the most recent messages
- ·Tier 3: sprint-status.yaml → find a story with status=
review - ·Tier 4: current git state — if branch ≠ main, confirm with the user
- ·Tier 5: ask the user directly
- ·Set
review_mode: full (spec/story file present) or no-spec
why- ·The cascade avoids asking the user unnecessarily — whatever is found automatically is used automatically
- ·
review_mode=full→ Acceptance Auditor runs;no-spec→ Acceptance Auditor is skipped
how- ·Construct the diff: staged / uncommitted / branch diff / commit range / provided diff — verify it's non-empty before continuing
- ·If the diff is > ~3000 lines → warn the user, offer chunking
- ·HALT + present a summary (diff stats, review_mode, loaded docs) for the user to confirm before reviewing
- ·
- 02Parallel review — 3 subagentsSpawn 3 adversarial reviewer subagents in parallel, each receiving different context.what
Dev spawns 3 parallel adversarial reviewer subagents, each receiving separate context:
- ·Blind Hunter: receives the diff ONLY — no spec, no project context. Finds logic errors, security issues, runtime exceptions. Doesn't know the intended behavior → catches unintended ones
- ·Edge Case Hunter: receives the diff + project read access. Traces all branching paths. Finds boundary errors, race conditions, null/empty/overflow cases
- ·Acceptance Auditor: receives the diff + spec file + context docs. Only runs when
review_mode=full. Checks whether each AC is implemented, AC drift, missing behaviors
why- ·The 3 perspectives are orthogonal — Blind Hunter: logic bugs; Edge Case: path gaps; Auditor: AC drift — no overlap, no misses
- ·A single reviewer misses at least one dimension — multi-agent review gives better coverage
how- ·All 3 run simultaneously (parallel), not sequentially
- ·If subagents aren't available → generate prompt files, HALT; the user runs each file in a separate session (ideally a different LLM) and pastes findings back
- ·If a layer fails or returns empty → log it into
failed_layers, continue with findings from the remaining layers
- ·
- 03Triage findingsNormalize → deduplicate → classify into 4 action buckets: decision_needed, patch, defer, dismiss.what
Automatically process findings from the 3 layers:
- ·Normalize: convert the 3 different output formats (Blind Hunter: markdown; Edge Case: JSON; Auditor: markdown) into a unified format: id, source, title, detail, location
- ·Deduplicate: an issue flagged by 2-3 reviewers is merged into one, with source set to e.g. "blind+edge"
- ·Classify each finding into exactly one bucket:
decision_needed(ambiguous, needs a human),patch(unambiguous fix),defer(pre-existing, not caused by this change),dismiss(noise/false positive) - ·Drop all
dismissfindings — log the count in the report
why- ·Without dedup, the human must sift through 3x the noise
- ·4 action buckets are clearer than severity labels — each bucket has a specific action
- ·Zero findings after triage → "Clean review" — without forcing a minimum number of findings
how- ·If
review_mode=no-spec: anydecision_neededfinding is reclassified →patch(if the fix is clear) ordefer - ·If
failed_layersis non-empty: warn the user before announcing the result
- ·
- 04Present, resolve & update statusRecord findings into the story file; HALT for the user to resolve decision_needed items and patches; update story status and sprint-status.yaml.what
- ·Record findings into the story file: append a "Review Findings" subsection under Tasks/Subtasks in the format
- [ ] [Review][Decision/Patch/Defer] title - ·Record
deferitems intodeferred-work.md - ·HALT: resolve
decision_neededfindings — the user decides (patch, defer, or dismiss) - ·HALT: process
patchfindings — the user chooses: apply all / leave as action items / walk through each - ·Update story status: all resolved →
done; action items remain →in-progress - ·Update sprint-status.yaml: story_key → new status, preserving ALL comments
why- ·Writing into the story file means findings persist in the source of truth — nothing is lost when the session ends
- ·
decision_neededfindings must have human input — AI cannot decide ambiguous choices on its own - ·Status
doneis set only once all issues are resolved — the story is never left in limbo
how- ·If
no-spec: no story file exists → present findings in chat only, don't persist them - ·If zero findings after triage → "Clean review" → skip straight to status update
- ·Suggest next: run Dev Story (next story) or re-run Code Review after fixes
- ·
- 01Inspect existing coverageRead current tests, coverage reports, and changed files.what
Read current tests, coverage reports, and changed files.
whyAutomation should fill real gaps, not duplicate existing evidence.
howMap gaps to risk priorities from Test Design.
- 02Select automation targetsChoose edge cases, regressions, integration points, and critical user journeys.what
Choose edge cases, regressions, integration points, and critical user journeys.
whyNot every path deserves E2E coverage.
howMatch test level to risk and maintenance cost.
- 03Generate testsAdd automated tests with realistic fixtures and assertions.what
Add automated tests with realistic fixtures and assertions.
whyUseful automation checks behavior and fails clearly.
howAvoid brittle selectors and over-mocked integration behavior.
- 04Run and stabilizeRun the suite repeatedly enough to catch flakiness.what
Run the suite repeatedly enough to catch flakiness.
whyFlaky tests erode trust in CI.
howFix timing, data, and isolation issues before handoff.
- 05Update CI notesDocument commands, test grouping, and CI configuration changes.what
Document commands, test grouping, and CI configuration changes.
whyAutomation only helps if the team can run it consistently.
howAdd config changes or explicit follow-up items.
- 01Load Context & Knowledge BaseDetermine review scope, detect stack, and load tiered knowledge base.what
Determine review scope (single / directory / suite), detect stack, load knowledge fragments, gather optional context (story file, test-design doc, framework config).
whyQuality evaluation rules are stack-conditional — frontend, backend, and fullstack have different patterns and anti-patterns.
howScan playwright.config.*, pom.xml, go.mod etc. Core knowledge always loaded; Playwright Utils and Pact.js Utils loaded on demand based on stack and config flags.
- 02Discover & Parse TestsCollect test files matching scope, parse metadata, and optionally capture browser evidence.what
Collect test files, parse metadata per file (framework, describe/test counts, fixtures, factories, network interception, waits, control flow), optionally capture browser trace and screenshots.
whyStructural analysis catches anti-patterns that static reading misses — flaky waits, shared mutable state, conditional logic inside tests.
howHalt if no tests found. Use Playwright CLI trace analysis for failing assertion root-cause when browser automation is enabled.
- 03Orchestrate Quality Evaluation (4 Workers)Run 4 parallel quality workers — Determinism, Isolation, Performance, Maintainability — then aggregate results.what
Dispatch all 4 quality workers in parallel; each writes a JSON result. Step 3F aggregates: weighted overall score, top-10 prioritised recommendations.
whyEach dimension requires different expertise — parallel workers are 60–70% faster than sequential and prevent one dimension from biasing another.
howResolve mode: agent-team → subagent → sequential (capability probe). Wait for all 4 outputs before aggregating.
- 04Generate Report & ValidateProduce test-review.md with overall score, dimension breakdown, critical findings, and recommended next workflow.what
Produce test-review.md: overall score + grade, dimension breakdown, critical findings with fixes, warnings, context references, and recommended next workflow.
whyA scored, graded report gives the team a concrete quality baseline and clear action items — not just a list of problems.
howValidate against checklist.md before finalising. Note: test-review scores quality, not coverage — direct coverage gaps to
bmad-testarch-trace.
- 01Load Context & Knowledge BaseCheck prerequisites, load knowledge base, and gather NFR thresholds from tech-spec, PRD, or story.what
Confirm implementation is accessible and evidence sources are available. Load knowledge base and context artifacts — NFR thresholds from tech-spec.md → PRD.md → story/test-design (priority order).
whyNFR evaluation without thresholds is guesswork — HALT if prerequisites are missing rather than proceeding with incomplete data.
howSummarise which sources were found and what evidence is available before moving to Step 2.
- 02Define NFR Categories & ThresholdsSelect the 8 standard ADR categories, extract thresholds, and build the NFR matrix.what
8 standard categories: Testability · Data Strategy · Scalability · Disaster Recovery · Security · Observability · Quality of Service · Deployability. Extract thresholds per category; mark UNKNOWN if not found.
whyAny UNKNOWN threshold means that category reports as CONCERNS — missing thresholds are never silently skipped.
howNever guess a threshold. Pull from sources in priority order: tech-spec → PRD → story/test-design.
- 03Gather EvidenceCollect measurable data per category — performance metrics, security reports, error history, DR drill records.what
Scan for: P95 response times, throughput, OWASP ZAP / Snyk reports, error rate history, burn-in results, DR drill records. Optionally capture browser network data via Playwright CLI.
whyMissing evidence is never "no issue" — any category without concrete evidence is automatically marked CONCERNS.
howStore browser capture outputs in
{test_artifacts}/nfr/. Close session after capture. - 04Orchestrate NFR Evaluation (4 Workers)Run 4 parallel domain workers — Security, Performance, Reliability, Scalability — then aggregate with worst-case rule.what
Dispatch Workers A–D in parallel: Security (04a), Performance (04b), Reliability (04c), Scalability (04d). Each outputs risk level (NONE / LOW / MEDIUM / HIGH) and PASS / CONCERN / FAIL per sub-category.
whyParallel workers are independent — one domain's risk does not bias another's evaluation.
howSub-step 4E aggregates all 4 outputs. Worst-case rule: overall risk = highest risk among 4 workers. Detect cross-domain compound risks (e.g. Performance + Scalability).
- 05Generate Report & ValidateProduce nfr-assessment.md with per-category results, overall risk level, remediation actions, and compliance matrix.what
Write nfr-assessment.md: per-category results with evidence summary, overall risk level, remediation actions (URGENT / HIGH / MEDIUM / LOW), compliance matrix (SOC2 / GDPR / HIPAA / PCI-DSS), CI gate YAML snippet.
whyA risk-rated report with remediation urgency lets the team prioritise fixes before the release gate — not after.
howValidate against checklist.md. Recommend next workflow: run
bmad-testarch-tracefor coverage gate, or proceed to release if NFR is PASS.
- 01Inventory requirementsList functional requirements, acceptance criteria, and NFRs.what
List functional requirements, acceptance criteria, and NFRs.
whyTraceability starts with a complete requirement set.
howMark ambiguous or changed requirements explicitly.
- 02Map tests to requirementsConnect each requirement to one or more test cases and result evidence.what
Connect each requirement to one or more test cases and result evidence.
whyCoverage must be visible at requirement level.
howInclude test file, case name, level, and latest result.
- 03Identify gapsFlag uncovered, weakly covered, flaky, or waived requirements.what
Flag uncovered, weakly covered, flaky, or waived requirements.
whyGaps are release risks and need owner decisions.
howClassify by severity and recommended action.
- 04Produce gate verdictSummarize PASS, CONCERNS, FAIL, or WAIVED.what
Summarize PASS, CONCERNS, FAIL, or WAIVED.
whyThe team needs a clear release posture.
howRequire rationale and owner for every concern or waiver.
- 05Hand off actionsSend required fixes to automation, review, Correct Course, or PM.what
Send required fixes to automation, review, Correct Course, or PM.
whyTraceability should trigger action, not just reporting.
howLink actions to requirement IDs and test evidence.
- 01InitializeConfirm change trigger; ask user description; verify access to PRD + Epics (required); choose mode: Incremental / Batch.what
Confirm change trigger; ask user description; verify access to PRD + Epics (required); choose mode: Incremental / Batch.
why· Trigger unclear → analysis misdirected; HALT early saves wasted effort
how· Ask numbered questions; if trigger still unclear after answer response → HALT and require specific details
- 02Analyze changeFollow checklist.md — 6 sections: trigger/context, epic impact, artifact conflicts, path forward evaluation, proposal components, final review. Record [x] / [N/A] / [!] each item.what
Follow checklist.md — 6 sections: trigger/context, epic impact, artifact conflicts, path forward evaluation, proposal components, final review. Record [x] / [N/A] / [!] each item.
why· Checklist systematic = does not miss impacts; ad-hoc analysis often misses downstream effects
how· Section by section, present progress after each major section
- 03Draft proposalsCreate explicit edit proposals per artifact — old → new format + rationale. Incremental: present each one; Batch: collect all, present together.what
Create explicit edit proposals per artifact — old → new format + rationale. Incremental: present each one; Batch: collect all, present together.
why· Old → new format = unambiguous, not misunderstanding about exactly what change
how· Story format: Story: [ID] Title / Section: ... / OLD: ... / NEW: ... / Rationale: ...
- 04Generate proposalCompile Sprint Change Proposal document: 5 sections → save to sprint-change-proposal-{date}.md . Present to user, ask Continue [c] / Edit [e].what
Compile Sprint Change Proposal document: 5 sections → save to sprint-change-proposal-{date}.md . Present to user, ask Continue [c] / Edit [e].
why· Structured document = stakeholder has specific review, approve, route that does not need re-explain context
how· Save to {planning_artifacts}/sprint-change-proposal-{date}.md
- 05Finalize & routeGet explicit user approval (yes/no/revise) → classify scope (Minor/Moderate/Major) → route: Minor→Dev / Moderate→PO+Dev / Major→PM+Architect. Update sprint-status.yaml.what
Get explicit user approval (yes/no/revise) → classify scope (Minor/Moderate/Major) → route: Minor→Dev / Moderate→PO+Dev / Major→PM+Architect. Update sprint-status.yaml.
why· Explicit approval = not ambiguity about "did we decide this?" — prevents team from acting on unapproved changes
how· Minor → Dev agent: deliverables = finalized edit proposals + implementation tasks
- 06CompleteSummary: issue addressed, scope classification, artifacts modified, routed to. Confirm deliverables and next steps.what
Summary: issue addressed, scope classification, artifacts modified, routed to. Confirm deliverables and next steps.
why· Explicit close = clear transition point; team knows workflow ended and who does what next
how· Personalized completion message to user; list every artifact already modified or created
- 01Locate fileFind sprint-status.yaml; if it does not exist → exit with guidance to run sprint-planning.what
Find sprint-status.yaml; if it does not exist → exit with guidance to run sprint-planning.
whyFind sprint-status.yaml; if it does not exist → exit with guidance to run sprint-planning.
howFind sprint-status.yaml; if it does not exist → exit with guidance to run sprint-planning.
- 02Parse & analyzeRead full sprint-status.yaml, classify keys (epics: epic-* / retros: *-retrospective / stories: remaining), count statuses (backlog/ready-for-dev/in-progress/review/done), detect risks (stale >7d, orphan story, in-progress epic no stories, story in review).what
Read full sprint-status.yaml, classify keys (epics: epic-* / retros: *-retrospective / stories: remaining), count statuses (backlog/ready-for-dev/in-progress/review/done), detect risks (stale >7d, orphan story, in-progress epic no stories, story in review).
whyRead full sprint-status.yaml, classify keys (epics: epic-* / retros: *-retrospective / stories: remaining), count statuses (backlog/ready-for-dev/in-progress/review/done), detect risks (stale >7d, orphan story, in-progress epic no stories, story in review).
howRead full sprint-status.yaml, classify keys (epics: epic-* / retros: *-retrospective / stories: remaining), count statuses (backlog/ready-for-dev/in-progress/review/done), detect risks (stale >7d, orphan story, in-progress epic no stories, story in review).
- 03Select recommendationPriority: in-progress story → review → ready-for-dev → backlog → retrospective → all done.what
Priority: in-progress story → review → ready-for-dev → backlog → retrospective → all done.
whyPriority: in-progress story → review → ready-for-dev → backlog → retrospective → all done.
howPriority: in-progress story → review → ready-for-dev → backlog → retrospective → all done.
- 04Display summaryProject info + story/epic counts + next recommendation + risk list.what
Project info + story/epic counts + next recommendation + risk list.
whyProject info + story/epic counts + next recommendation + risk list.
howProject info + story/epic counts + next recommendation + risk list.
- 05Offer actions[1] run recommended workflow / [2] show stories by status / [3] show raw yaml / [4] exit.what
[1] run recommended workflow / [2] show stories by status / [3] show raw yaml / [4] exit.
why[1] run recommended workflow / [2] show stories by status / [3] show raw yaml / [4] exit.
how[1] run recommended workflow / [2] show stories by status / [3] show raw yaml / [4] exit.
- 01Epic discoveryFind the recently completed epic from sprint-status; confirm with the user; verify all stories are done.what
Find the recently completed epic from sprint-status; confirm with the user; verify all stories are done.
why· Retro on epic not yet done = missing data → insights insufficient
how· If partial: the user chooses whether to continue or HALT to finish stories first
- 02Deep story analysisRead EACH story file in the epic — extract dev notes, review feedback, lessons, tech debt, testing insights → synthesize cross-story patterns.what
Read EACH story file in the epic — extract dev notes, review feedback, lessons, tech debt, testing insights → synthesize cross-story patterns.
why· Story files contain real data — rich material for discussion, not memory-based assumptions
how· Read each file: {implementation_artifacts}/{N}-{M}-*.md
- 03Previous retro integrationLoad retro epic before — check action item follow-through, lessons were applied or skipped.what
Load retro epic before — check action item follow-through, lessons were applied or skipped.
why· Not checking the previous retro = repeating the same mistakes and losing accountability
how· If this is epic 1 (no previous retro): set first_retrospective = true and skip this step
- 04Next epic previewLoad epic {N+1} — identify dependencies, gaps, preparation needs before start.what
Load epic {N+1} — identify dependencies, gaps, preparation needs before start.
why· Prep work identified during the retro → start epic {N+1} with solid foundation instead of hit blockers mid-epic
how· If epic {N+1} is not yet defined: skip, but still proceed with the retro normally
- 05Epic review discussionParty mode: Amelia facilitates what went well / what didn't / patterns — INTERACTIVE with user, blameless, system-focused.what
Party mode: Amelia facilitates what went well / what didn't / patterns — INTERACTIVE with user, blameless, system-focused.
why· Party mode = psychological safety — everyone speaks honestly, not blame, not defensive
how· Ground rules: no blame, focus on systems, specific examples preferred, every voice counts
- 06Next epic prep + action itemsSynthesize action items; plan next epic preparation tasks; significant change detection; critical readiness check.what
Synthesize action items; plan next epic preparation tasks; significant change detection; critical readiness check.
why· Significant change detection prevents starting epic {N+1} on wrong assumptions → mid-epic failure
how· Detect significant change: architectural assumptions wrong, scope change, tech approach must change, newly discovered dependencies, security/compliance issues
- 07Save & update statusSave epic-{N}-retro-{date}.md ; update sprint-status: epic-{N}-retrospective → done .what
Save epic-{N}-retro-{date}.md ; update sprint-status: epic-{N}-retrospective → done .
why· Date-stamped file = append-only institutional memory — next retro load file this to check follow-through
how· If the epic-{N}-retrospective key is not found in sprint-status → warn the user; manual update is needed
- 01Clarify and routeClarify the requested change and route to the appropriate execution path.
- 02PlanProduce an implementation plan for the clarified change.
- 03ImplementImplement the planned change.
- 04ReviewReview the implementation and finalize the iteration.