B
BMAD METHOD Breakthrough Method for AI-Driven Agile Development + WDS · TEA
v6.10.0 BMM + WDS + TEA Unified Process Map

BMAD-METHODv6.10.0

BMM + WDS + TEA
BMM · WDS · TEA — one unified process map
Click section heading to collapse / expand · click "steps" to see step details · click "what · why · how" for the full skill guide
BMM Overview
BMM · Core Method
06
BMM Agents
29+
BMM Skills
16+
BMM Workflows
04
BMM Phases
06
BMM Skill Groups
WDS · Whiteport Design Studio
03
WDS Agents
09
WDS Skills / Workflows
09
WDS Phases
12
WDS Guide Files
25+
Yrs UX Foundation
TEA · Test Architect Enterprise
01
TEA Agent
09
TEA Skills / Workflows
40+
Knowledge Fragments
7
Test Levels
5
CI Platforms
Agent Team
BMM v6.10.0 Agents
6 core BMM agents · Each owns a role · Collaborates through workflows
6 AGENTS
6 BMM
01
A
Analyst
Mary · Discovery
Explores the problem space, facilitates brainstorming, conducts market/domain/technical research, and writes strategic product briefs.
01·A01·B01·C01·DDP
02
P
PM
John · Product Manager
Turns vision into executable requirements — PRD (single intent-detecting skill: create/update/validate), Epics & Stories, Readiness Gate, and Course Correction.
02·A03·B03·C04·I
03
X
Architect
Winston · System Design
Makes well-reasoned technical decisions (ADRs) — tech stack, data, API, security, deployment, and NFR validation. Partners with PM on the Implementation Readiness Check.
03·A03·B03·C
04
D
Dev
Amelia · v6.3+ consolidated
v6.3+ unified — Amelia consolidates Bob (SM), Quinn (QA), and Barry (solo-dev).
04·A04·C04·I04·SS04·J
05
U
UX Designer
Sally · Experience Design
Designs the end-to-end experience — user journeys, IA, wireframes, interaction patterns, and WCAG 2.1 AA accessibility.
02·BCU
06
W
Tech Writer
Paige · Documentation
BMM core agent — documents existing projects (DP), writes documentation (WD), updates standards (US), creates Mermaid diagrams (MG), validates documentation (VD), and explains concepts (EC). Conversational triggers WD/US/MG/VD/EC require a description argument.
DPWDUSMGVDEC
Skills
29 Skills in BMM v6.10.0
Each skill is a discrete capability · Agents wrap menus of skills · Many skills are shared across agents
29 SKILLS
7 groups · shared
🔗 SHARED SKILLS — Available to all agents — not owned by any specific agent 14 skills
bmad-help UPDATED v6.10.0
Universal routing assistant — auto-invoked at the last step of every workflow. Checks project state and recommends the most appropriate next workflow.
DRIVESLast step of all workflowsI/Oproject state → next-workflow recommendationTECHNIQUESState InspectionRoadmap HeuristicsGUARDRAILOnly recommends — never auto-runs the next workflow.
bmad-advanced-elicitation UPDATED v6.10.0
Optional after each step — stress-tests assumptions, validates against similar products, explores alternatives. Adds Subtraction (counters additive bias) and Map Is Not the Territory (guards against over-trusting a lossy model). Menu [A] appears at the end of each micro-file.
DRIVESOptional [A] in 01·C/D, 02·A/B, 03·AI/Ocurrent draft → challenged assumptions + alternativesTECHNIQUESAssumption Stress-testDevil’s AdvocateWhat-ifSubtractionMap Is Not the TerritoryGUARDRAILFires only on explicit [A]; never redirects unprompted.
bmad-party-mode UPDATED v6.10.0
Multi-agent collaborative session — reborn with creatable/savable custom parties, a built-in anti-consensus club (Wildcard, Level, Killjoy, Splinter via --party=anti-concensus-club), optional persistent per-party memory (memlog), and pacing/dynamics rules. Open-ended by default — ends only on an explicit signal, with an opt-in --non-interactive flag; standing agents are kept alive and resumed. 4 run modes: session / auto / subagent / agent-team (point-to-point relay, not broadcast). Ships with a preloaded "Code Review Crew" party.
DRIVESTradeoff reviews · Retrospective · edge-case designI/Oshared context → cross-agent debate + decisionTECHNIQUESRound-robin FacilitationRoster LookupCustom Party AuthoringGUARDRAILNeeds party-mode visibility in config; each agent keeps its persona.
bmad-party-mode
v6.9 updated — configurable custom personas, named rooms with optional scenes, four run modes (session/auto/subagent/agent-team), persistent session memory under {memory_dir}/<party_id>/.
DRIVESTradeoff reviews · Retrospective · edge-case design · 7 activation stepsI/Oshared context → cross-agent debate + decision + party memoryTECHNIQUESRound-robin FacilitationRoster LookupCustom Party AuthoringGUARDRAILEach agent keeps its persona; persistent memory per party_id.
bmad-forge-idea UPDATED v6.10.0
Persona-driven interrogation that pressure-tests an idea through open-ended socratic questioning until it hardens, proves out, or dies cheaply. Usable at any phase, by any agent or standalone. Optional handoff to bmad-spec or bmad-quick-dev.
DRIVESanytime — no fixed phaseI/Oraw idea → forged-idea.md (optional) + forge-report.htmlTECHNIQUESSocratic InterrogationPersona-driven Pressure-testGUARDRAILOpen-ended — no fixed step count; lets a weak idea die cheaply rather than propping it up.
bmad-brainstorming
Core skill for 01·A. Library of ~60 ideation techniques plus a visual technique composer (brain-selector.html, 108 techniques), session orchestration, clustering and prioritization. 3 selectable stances — Facilitator / Creative Partner / Ideate-for-me. Append-only memlog with optional --by user/--by coach authorship attribution.
DRIVES01·A (6 steps)I/Otopic + goal → brainstorming-report.mdTECHNIQUESSCAMPERSix HatsReverseImpact/Effort 2×2MoSCoWGUARDRAILNever ideates for the user; pivot every 10 ideas to break anchoring.
bmad-spec
Distills any intent input — brain dump, PRD, GDD, RFC, brief, transcript — into SPEC.md, the five-field kernel (Why, Capabilities, Constraints, Non-goals, Success signal) plus companion files for load-bearing content that doesn't fit the kernel. The canonical machine contract every downstream BMad skill consumes.
DRIVESAny phase — create / update / validateI/Oany intent input → SPEC.md + companions + .memlog.mdTECHNIQUESLoad-bearing TestPreservation SweepSpec Law (8 rules)GUARDRAILNever invents an answer for a recognized gap — logs it as an open question instead.
bmad-customize
Authors and updates customization overrides for installed BMad agents and workflows — translates plain-English intent into a correctly-placed TOML override under _bmad/custom/ (team-committed or user-gitignored), then verifies it resolves.
DRIVESOn demand — any agent or workflow skillI/Oplain-English change request → *.toml / *.user.toml overrideTECHNIQUESSurface ClassificationTeam/User MergeGUARDRAILNever silently overwrites an existing override; always shows a diff and waits for explicit yes.
bmad-shard-doc
Splits a large markdown document into smaller, organized files based on level-2 sections using @kayvan/markdown-tree-parser — generates an index.md and offers to delete, archive, or keep the original.
DRIVESOn demand — any large markdown docI/Olarge .md file → destination folder of section files + index.mdTECHNIQUESLevel-2 ShardingGUARDRAILHalts if no files are produced; never leaves both original and shards without asking.
bmad-index-docs
Generates or updates an index.md that references every file in a target folder, grouped by type/subdirectory with content-derived (not filename-derived) descriptions.
DRIVESOn demand — any target folderI/Otarget folder → index.mdTECHNIQUESContent-derived DescriptionsGUARDRAILReads each file rather than guessing from its filename.
bmad-editorial-review-prose
Clinical copy-editor pass — reviews text for communication issues that impede comprehension (Microsoft Writing Style Guide baseline) and outputs a three-column Original/Revised/Changes table. Never challenges ideas, only how they're expressed.
DRIVESOn demand — any prose reviewI/Ocontent (+ optional style guide) → 3-column fix tableTECHNIQUESMinimal InterventionReader-type CalibrationGUARDRAILCONTENT IS SACROSANCT — never challenges ideas, only clarifies expression.
bmad-editorial-review-structure
Structural editor — run before copy editing. Maps document structure against the right model (Tutorial, Reference, Explanation, Prompt/Task, Strategic/Pyramid) and proposes CUT / MERGE / MOVE / CONDENSE / QUESTION / PRESERVE recommendations, never executes them.
DRIVESOn demand — before copy editingI/Odocument (+ purpose/audience) → prioritized recommendation listTECHNIQUESStructure ModelsFront-load ValueGUARDRAILPropose, don't execute — user decides what to accept.
bmad-review-adversarial-general
Cynical, jaded critical review of any artifact — diff, spec, story, doc — with zero patience for sloppy work. Finds at least ten issues and outputs a plain Markdown findings list.
DRIVESOn demand — any artifact needing a critical passI/Ocontent (+ optional focus areas) → Markdown findings listTECHNIQUESAdversarial SkepticismGUARDRAILZero findings is suspicious — re-analyze rather than pass silently.
bmad-review-edge-case-hunter UPDATED v6.10.0
Pure path tracer — mechanically walks every branch and boundary condition in a diff, file, or function and reports only unhandled paths as a strict JSON array. Now generalizes to named-set/enum/status-code/sentinel/flag partial-handling gaps, and folds the deletion-contract audit in as a gated pass inside this same turn. Orthogonal to adversarial review: method-driven, not attitude-driven.
DRIVESOn demand — diffs, files, functionsI/Odiff / file / function → JSON array of unhandled pathsTECHNIQUESExhaustive Path EnumerationGUARDRAILReports only unhandled paths — never editorializes on code quality.
📊 Mary — Business Analyst 5 skills
bmad-market-research
Market analysis — TAM/SAM/SOM, competitor mapping, trends, user sentiment, pricing benchmarks. Triangulates data from web, official docs, and expert content.
DRIVES01·B (market track)I/Oresearch questions → market findingsTECHNIQUESTAM/SAM/SOMCompetitor MappingPESTELTriangulationGUARDRAILEvery figure needs a URL; ≥3 sources before claiming a pattern.
bmad-editorial-review-structure
Structure quality review — evaluates section order, information hierarchy, completeness against template, and logical flow. Returns a structured gap report.
DRIVESanytime · any artifact documentI/Odraft document → structure-review.md with gap reportTECHNIQUESTemplate ConformanceHierarchy AuditGUARDRAILVerify against the canonical template before flagging as missing.
bmad-index-docs
Builds and maintains a navigable index of all project documents — auto-detects artifact types, generates a linked table of contents, keeps it up to date across workflow runs.
DRIVESanytime · post-workflow index updateI/Odocs/ folder → docs/INDEX.mdTECHNIQUESArtifact DetectionTOC GenerationGUARDRAILIndex only — never modifies the source documents.
bmad-review-adversarial-general
Adversarial review of any artifact — attacks assumptions, finds logical gaps, stress-tests requirements, identifies unstated dependencies. Red-team framing.
DRIVESanytime · gate before handoffI/Oartifact → adversarial-review.md with ranked findingsTECHNIQUESRed-teamAssumption AttackGap AnalysisGUARDRAILUse a different LLM to avoid self-confirmation bias.
bmad-review-edge-case-hunter
Systematically hunts edge cases in requirements, architecture, or code — boundary conditions, error paths, concurrency issues, missing invariants. Outputs a prioritized edge-case register.
DRIVESanytime · pre-implementation gateI/Oartifact → edge-case-register.md ranked by riskTECHNIQUESBoundary ValueError Path MappingInvariant CheckGUARDRAILRank by impact — don’t let low-risk edge cases inflate scope.
bmad-shard-doc
Splits a large document into context-sized shards for multi-agent workflows — adds shard headers, cross-reference links, and a manifest so downstream agents can reconstruct the full picture.
DRIVESanytime · large document pre-processingI/Olarge doc → shards/ + shard-manifest.mdTECHNIQUESSemantic ChunkingCross-reference LinksGUARDRAILNever lose content — manifest must account for every section of the original.
bmad-spec
Generates a formal specification document for a given component or interface — precise inputs/outputs, invariants, constraints, and acceptance criteria in a machine-readable + human-readable format.
DRIVESanytime · pre-implementation spec gateI/Ocomponent description → spec.md with invariants + ACTECHNIQUESContract-firstInvariant DeclarationGherkin ACGUARDRAILSpec must be falsifiable — every claim needs a test that could fail it.
📋 John — Product Manager 4 skills
bmad-prd
Core skill for 02·A — single intent-detecting workflow: Create / Update / Validate, with .memlog.md audit trail, addendum.md overflow doc, Essential Spine + Adapt-In Menu, and a Reviewer Gate with HTML report. Supersedes the deprecated create/edit/validate PRD shims. Scale-adaptive depth Level 0–4 × domain complexity Low/Medium/High.
DRIVES02·A (13 steps, facilitator)I/Oproduct-brief + research → PRD.mdTECHNIQUESFR/NFR TaxonomyUser StoriesGherkin ACSMARTPersonasGUARDRAILEach FR traces to a validated need; each NFR has a measurable threshold.
bmad-create-epics-stories
Orchestrates 03·B — builds FR ↔ component traceability matrix (cross-referencing architecture), INVEST validation, dependency graph cycle detection. Shared with Winston.
DRIVES03·B (PM + Architect)I/OPRD + ARCHITECTURE-SPINE.md → epic-{N}.md + stories/*.mdTECHNIQUESINVESTStory SlicingDependency GraphTraceabilityGUARDRAILEach story traces to an FR; the dependency graph must be acyclic.
bmad-implementation-readiness
Cross-cutting validation shared with PO — cohesion check matrix: PRD ↔ UX ↔ Architecture ↔ Stories. Last gate before Phase 4.
DRIVES03·C (gate before Phase 4)I/OPRD + architecture + epics → readiness-report.mdTECHNIQUESCohesion CheckTraceability MatrixGate ReviewGUARDRAILIndependent validation, ideally a different LLM.
bmad-correct-course
Mid-sprint pivot management — variance analysis + gap analysis + blast radius. 3-tier response: Patch / Re-route / Escalate based on impact scope.
DRIVES04·I (5 steps)I/Osprint-status + deviation → updated planTECHNIQUESVariance AnalysisBlast-radiusPatch/Re-route/EscalateGUARDRAILEscalate when blast-radius exceeds the sprint; no patches hiding structural issues.
🎨 Sally — UX Designer 1 skills
bmad-create-ux-design
Core skill for 02·B — orchestrates 10 steps: journey mapping → IA → wireframing → interaction patterns → component specs → responsive → a11y → design tokens → finalize.
DRIVES02·B (10 steps)I/OPRD.md → DESIGN.md + EXPERIENCE.mdTECHNIQUESJourney MappingIAWireframingWCAG 2.1 AADesign SystemsGUARDRAILa11y is first-class; every screen gets edge/empty/error states.
🏗️ Winston — System Architect 1 skills
bmad-architecture
Core skill for 03·A — intent-routed (create / update / validate) lean-spine model: fixes only the invariants (paradigm, boundaries, ownership) a future builder can't read off compliant code, produces ARCHITECTURE-SPINE.md as source of truth, with a breadth-coverage rubric and a Reviewer Gate. Supersedes the deprecated fixed 11-step bmad-create-architecture flow.
DRIVES03·A (8 steps)I/OPRD + DESIGN.md/EXPERIENCE.md / SPEC.md → ARCHITECTURE-SPINE.md (+ ADRs, optional renderings)TECHNIQUESCoaching/Fast PathReviewer GateADR (Nygard)Breadth-coverage RubricGUARDRAILFix an invariant only when two independently-built units could otherwise diverge; everything else is seed or Deferred.
💻 Amelia — Developer Agent (consolidated v6.3+) 7 skills
bmad-sprint-planning
One-time per project — loads all epics, sequences stories via topological sort + WSJF, groups into sprints with capacity planning, generates sprint-status.yaml as the single source of truth.
DRIVES04·A (once per project)I/Oall epics → sprint-status.yamlTECHNIQUESTopological SortWSJFSprint Boundary heuristicsGUARDRAILRespect dependency order; don’t oversize a sprint.
bmad-create-story
Just-in-time story prep — expands epic-level AC into executable spec, extracts technical context from architecture, defines test strategy and Definition of Done.
DRIVES04·C (just-in-time)I/Onext epic AC → story-[slug].mdTECHNIQUESAC ExpansionINVESTDefinition of DoneTest StrategyGUARDRAILPull technical context from architecture, don’t invent it.
bmad-dev-story
TDD-disciplined implementation — activates the skipped tests from TEA's ATDD scaffold, observes red, implements just enough to go green, then validates the full suite and updates story status. Never claims tests pass without running them.
DRIVESDev Story (per story)I/Ostory file + ATDD scaffold → code + tests + updated statusTECHNIQUESTDDRed-Green-RefactorDefinition of DoneGUARDRAILTests before code; implement just enough to pass — no gold-plating.
bmad-code-review UPDATED v6.10.0
Adversarial code review — 3 parallel subagents (Blind Hunter, Edge Case Hunter, Acceptance Auditor) triage findings into patch / defer / decision-needed. Severity calibration now requires reading surrounding source (call sites, guards) before rating — the old "prefer conservative when uncertain" tie-breaker is gone — and deletion-contract checks are delegated to bmad-review-edge-case-hunter.
DRIVESCode Review (per story, after Dev Story)I/Odiff + tests → review findings + patches/deferralsTECHNIQUESReview triageFeedback receptionAcceptance auditGUARDRAILResolve review findings before TEA evidence gates (Automate, NFR, Traceability).
bmad-qa-generate-e2e-tests
E2E test generation workflow — reads the story like a tester: scores each AC scenario by risk, maps it to test cases (happy path · edge cases · error paths), then generates complete Playwright spec files with fixtures, page models, and data factories. Five sequential steps: load story → risk scoring → test plan → generate spec → validate & wrap up.
DRIVES04·E / 04·F (TEA post-code evidence)I/Ostory-[slug].md + AC → tests/e2e/[feature].spec.tsTECHNIQUESRisk ScoringAC-to-Scenario MappingFixture-first DesignSelector ResilienceGUARDRAILScope is the AC — do not generate tests for behavior outside acceptance criteria. Prefer data-testid; each test must run independently and must not depend on order.
bmad-retrospective
Per-epic retrospective — data-based scorecard, what went well (appreciative inquiry), friction points (blameless framing), root cause analysis (5 Whys), max 3 SMART action items.
DRIVES04·J (per epic)I/Ocompleted stories → retro-{epic}.mdTECHNIQUESAppreciative InquiryStart/Stop/Continue5 WhysGUARDRAILEvery action item gets a named owner, not "the team".
bmad-dev-auto NEW v6.10.0
One iteration of an unattended development loop — clarify & route, plan, implement, review. Invoked by name (not yet menu-chained).
DRIVESPhase 4 — invoked by nameI/Ospec file → implementation_artifacts (spec file + bmad-dev-auto-result-*.md)TECHNIQUESClarify & RoutePlanImplementReviewGUARDRAILOne iteration per invocation — does not chain further iterations on its own.
✍️ Paige — Technical Writer 2 skills
bmad-document-project
Documents an existing project from scratch — codebase analysis, doc structure, initial documentation generation. Triggered by DP.
DRIVESDP triggerI/Ocodebase → docs/ structure + initial docsTECHNIQUESDiataxisCodebase AnalysisGUARDRAILAsk "who reads this?" first — no documentation theater.
bmad-agent-tech-writer
Consolidated tech-writer skill — 5 intents: write (WD) · update-standards (US) · validate (VD) · explain (EC) · mermaid diagram generation. Replaces standalone bmad-write-document, bmad-update-standards, bmad-validate-doc, bmad-explain-concept, bmad-mermaid-generate.
DRIVESWD · US · VD · EC · mermaid triggers · anytimeI/Ointent + args → targeted doc / diagram / standards / validationTECHNIQUESDiataxisMermaid DiagramsAudience-tieredGUARDRAILEach intent stays in its Diataxis lane; use a different LLM for validate intent.
WDS
Whiteport Design Studio v0.4.3
Design-led add-on module · 3 specialist agents (Saga · Freya · Mimir) · Spec-driven pipeline from discovery to working code
Add-on
What is WDS?

A free, open-source design workflow giving designers expert AI agents to guide them through strategy and design process — making impactful deliveries for AI-driven or traditional team development.

WDS unites entrepreneurs, developers, and designers under a shared language for the AI era. Under the hood: a structured, AI-assisted pipeline taking a product from raw idea to production-ready implementation, coordinated by three specialized AI agents.

BMAD-native: runs via the BMAD Method Module — compatible with Claude Code, Cursor, Windsurf, Codex, and 16+ IDEs. MIT license · free · community-maintained.

Core Philosophy
Designers are needed more than ever
Designers guide decisions in AI-driven development. Your expertise in dialogues with users and stakeholders is invaluable — especially when AI amplifies your impact.
Brilliant design is not a coincidence
Great design happens when strategy, user needs, and business goals align. WDS agents amplify designer skills and experience — they don't drown the process.
The specification is the new code
AI generates code but struggles to change it. Your design documentation becomes the source of truth — replacing months of prompting with strategic specs that guide development.

A method that unites

WDS is an AI Agent framework that creates a unified language for entrepreneurs, developers, and designers to collaborate in the AI era. WDS assists you in strategic design leadership for any digital product from idea to polished product.

Your design replaces prompting

You map user psychology and business goals. You envision the user interaction and create conceptual specifications that guide development. You transform vision into strategic design thinking. AI amplifies your expertise. With WDS, you're the strategic center that holds it all together.

Powered by BMad Method

WDS is a module for designers within the BMad Method. Instead of mindless prompting, the BMad method is most effective when run in an IDE like Cursor, VS Code, Windsurf, or similar tools.

One chat leads to the next

The method preserves AI capabilities by dividing creative conversations into dialogues that result in high-level documents. These documents become the perfect prompts for the next step — preserving the AI's context window.

Built for collaboration

Everything is saved and published using GitHub — the same technology developers use. By delivering design in the perfect form for development, designers can deliver far more value than ever before.

Conceptual Specifications For designers that mean business
Experimentation is essential

Napkin sketches, whiteboard drawings, Figma prototypes, code snippets, inspiration boards — all of it. This is where creativity happens. The playful creative process reaches a point when experimentation leads to something great, something we want to build and bring to the world.

Your designs are the perfect prompt

When you gather all your experiments in one place, defining them as user scenarios with step-by-step interaction and clear explanations — not just what you designed, but why — you give AI all it needs to make the final product reality.

The specification becomes the new code

In the world of AI development, the specifications become the product. The code is just output — generated and regenerated from your design work. Your specifications are the source of truth that gets maintained and refined over time.

Design once, iterate forever

Want to change something? Update the specifications. AI regenerates the code — clean, consistent, and aligned with your design intent. No spaghetti code. No lost reasoning. No painful refactoring. Your design work IS the product.

Endorsed by

"Whiteport Design Studio adds a strategic design dimension that strengthens BMad: a practical way to translate intent into interfaces and product direction. It fits the spirit of the agent framework, with expert-driven workflows where AI is a partner, not an autopilot."

B
BMadCode
Founder, BMad Method
Core Properties
Design in the center
Your strategic design thinking drives every step — WDS amplifies expertise, not replaces it
Structured pipeline
9 phases (0–8), each with defined inputs, outputs, and quality gates — no ad-hoc guessing
Agent-driven, designer-led
Saga, Freya, and Mimir amplify your skills without drowning the creative process
Brownfield capable
Phase 8 (Product Evolution) handles existing products without redoing everything
9-Phase Pipeline · Phase 0 optional
0 opt.
Alignment & Signoff
Secure formal stakeholder alignment before any design or development begins. Produces alignment pitch, formal contract, and internal signoff. Optional — required for client/consultant scenarios.
Saga
1
Project Brief
The North Star. Establishes the complete strategic foundation every downstream phase reads — vision, positioning, business model, target users, visual direction, platform requirements. 36-step facilitated flow with Workshop / Suggest / Dream modes.
Saga
2
Trigger Mapping
Connects business goals to user psychology via Effect Mapping — maps positive and negative driving forces (aspirations, fears, blockers) behind behavior, then surfaces which features have the most impact. Produces Trigger Map, persona files, and Feature Impact Matrix with Mermaid diagrams.
Saga
3
UX Scenarios
Transforms the Trigger Map into concrete user journey outlines — linear "sunshine paths" that expose every page and decision point to design in Phase 4. One scenario per user goal; each outline includes persona, entry/exit states, step-by-step journey, pages required, and SEO considerations.
Freya
4
UX Design
The core design phase — "the spec is the new code." Transforms scenarios into development-ready written specifications: content, interaction logic, spacing, typography, state variations, and object registry per page. Activities: Conceptualize [C] · Sketch [K] · Suggest [S] · Dream [D] · Specify [P] · Validate [V] · Visual Design [W] · Design System [M] · Design Delivery [H].
Freya
5
Agentic Development
Spec-driven development. Mimir reads Freya's Work Order completely, writes a PRD, then implements one verified requirement at a time with browser testing. Activities: Prototyping [P] · Development [D] · Bugfixing [F] · Evolution [E] · Analysis [A] · Reverse Engineering [R] · Acceptance Testing [T].
Mimir
6
Asset Generation
Where specs become pixels. Generates production-ready visual assets from specifications using Stitch and Nano Banana — wireframes, page designs, UI elements, icon sets, photography, illustrations, video/motion, and written copy. Activities: Wireframes [W] · Page Designs [P] · UI Elements [U] · Icons [I] · Images [M] · Videos [V] · Content [C] · Export to Figma [E].
Freya
7
Design System optional
Builds a living component library and design token system extracted from actual page specs — grows organically from real patterns, not designed upfront. Produces component spec files, design-tokens.md, and component-library-config.md. Activities: Create [C] · Import [I] · View [V] · Edit [E] · Browse [B].
Freya
8
Product Evolution
Brownfield improvements — the full WDS pipeline compressed into a tight kaizen loop for living products. One view, one scenario, one improvement per branch-based cycle. Activities: Analyze [A] · Scope [S] · Design [D] · Implement [I] · Test [T] · Deploy [P]. Can feed back into Phases 2, 3, and 4 when new user types emerge.
Freya
Mimir
WDS Embedded

Saga drives strategy in analysis (alignment · project brief · trigger mapping). Freya drives design in planning — PRD → Phase 3 scenarios → Phase 4 (UX design · conceptual specs · spec audit) → Phase 6 assets → Phase 7 design system (optional). After Phase 7, two workflows run in parallel: Design Delivery (Phase 4, Freya — packages flows for development handoff) and Create UI Prototype (Phase 5, Mimir — validates design in code). Mimir then drives the build in implementation (tech audit → PRD → build, one verified requirement at a time), reading the WDS specs — governed by BMAD Dev Story discipline and gated by Murat (TEA). Freya's Phase 8 runs as standalone brownfield support.

Implementation — Mimir-led Split

Mimir owns planning (tech-audit → write-prd, derived straight from the WDS page specs so design intent is preserved — something a plain BMAD Dev Story can't read). BMAD Dev Story then covers execution for anything outside Mimir's build loop: TDD implementation → systematic debugging → code review, run by Amelia. Dev Story's own task planning is intentionally deferred to Mimir for WDS-enabled stories — write-prd is the single planning authority, avoiding two systems owning the task list. The build/story overlap is not a conflict: Mimir's "one requirement at a time" is the rule, and Dev Story's TDD discipline is the mechanism that realizes it.

WDS Specialist Team · 3 Agents
W1
S
WDS Analyst
Saga 📚 · Strategy & Discovery
WDS Phases 0–2 — replaces Mary on WDS projects. Alignment & Signoff, Project Brief (the North Star), and Trigger Mapping (Effect Mapping — business goals → user psychology). Produces WDS-native artifacts Freya consumes directly.
ASPBTM
W2
F
WDS Designer
Freya 🎨 · Strategic UX
WDS Phases 3–8 — replaces Sally on WDS projects. Scenarios, UX Design (the spec is the new code), Asset Generation, Design System, Design Delivery, and brownfield Product Evolution. Writes development-ready specs Mimir implements.
SCUXSPSAGADSDDPE
W3
M
WDS Builder
Mimir 🔨 · Spec-Driven Dev
WDS Phase 5 · net-new — no prior BMM equivalent. Reads Freya's Work Order, runs the tech audit, writes the PRD, and builds one verified requirement at a time with browser testing. Strongest paired with Amelia's BMAD Dev Story governing execution outside the WDS build loop.
TAPRBU
🧪
BMAD TEA
Test Architect Enterprise · Add-on Module · Agent Murat
Add-on

Test Architect Enterprisev1.19.0

Advanced module · Add-on for BMAD core · Murat is your guide
9
Skills / Workflows
40+
Knowledge Fragments
7
Test Levels Supported
5
CI Platforms
1
Agent · Murat 🧪
I
What is BMAD TEA?
Test Architect Enterprise — advanced test strategy module
Add-on Module
Separate Installation

BMAD TEA (Test Architect Enterprise) is a standalone add-on module installed on top of BMAD core to extend comprehensive test architecture capabilities. While BMAD core provides bmad-qa-generate-e2e-tests for quick test needs within a sprint, TEA elevates to a different strategic tier — from risk assessment to CI/CD governance, from contract testing to NFR evidence audit.

TEA does not replace BMAD core — TEA complements core. All TEA workflows integrate with BMAD story/PRD/epics artifacts and use the same config pattern (TOML customization layer).

Comparison Dimension BMAD Core BMAD TEA
InstallationIncludedSeparate add-on
Test agentAmelia (QA role)Murat 🧪 (specialized)
Test strategyHappy path + critical errorsRisk-based, full coverage matrix
Test levelsAPI + E2EUnit + Component + API + E2E + Contract + NFR
CI/CDRun test commandCI pipeline scaffold + quality gates
Knowledge base—40+ fragments, JIT loading
Release gate—Trace coverage + NFR evidence audit
II
Murat — Master Test Architect
BMAD TEA's dedicated agent · 🧪
Trigger: "talk to Murat"
or "Test Architect"
Identity
Murat
Master Test Architect & Quality Advisor · 🧪
Expert in risk-based test architecture. Equally proficient in pure API testing (pytest, JUnit, Go test, xUnit, RSpec) and browser E2E (Playwright, Cypress), contract testing (Pact), and performance testing (k6).
Communication Style

"Strong opinions, weakly held" — Murat's mantra. Always willing to change perspective when new data emerges.

Speaks in risk calculations and impact assessments. Combines data with intuition — no purely theoretical judgments, always grounded in project reality.

Supported Platforms
Test frameworks
Playwright · Cypress · pytest · JUnit · Go test · xUnit · RSpec · k6
CI/CD platforms
GitHub Actions · GitLab CI · Jenkins · Azure DevOps · Harness CI
§
7 Core Principles
Murat makes decisions based on these 7 principles
01
Risk-based testing — Testing depth is proportional to risk level. Not everything needs the same level of testing.
02
Data-driven quality gates — No subjective judgments. Every quality gate must have measurable evidence.
03
Tests reflect usage patterns — Whether API, UI, or both — tests must mirror how users actually use the system.
04
Flakiness is serious technical debt — Flaky tests are not "slightly annoying" — they are accumulated debt that erodes confidence in the entire test suite.
05
Calculate risk vs value — Every testing decision must pass the question: "Does the cost of this test justify the protection value it provides?"
06
Prefer lower test levels — Unit > Integration > E2E. Lower test levels run faster, are more stable, easier to debug — always use the lowest level that can verify the behavior.
07
API tests are first-class citizens — Not accessories to UI tests. API tests are faster, more stable, and directly cover business logic — they deserve dedicated investment.
§
9 Skills Murat Dispatch
Invoke by code or natural description — Murat auto-routes to the correct skill
Code Skill Description
TMTbmad-teach-me-testingInteractive testing academy — 7 sessions from fundamentals to advanced, with quizzes and role-based learning paths.
TDbmad-testarch-test-designDesign test strategy — risk assessment, NFR planning, coverage strategy for a system or epic.
TFbmad-testarch-frameworkScaffold a production-ready test framework — Playwright or Cypress, with fixtures, helpers, and full configuration.
CIbmad-testarch-ciAdvise and scaffold a CI/CD quality pipeline — quality gates, sharding, parallelization, reporting.
ATbmad-testarch-atddATDD — generate failing acceptance tests + implementation checklist before dev writes any code.
TAbmad-testarch-automateExpand coverage — generate API/E2E tests, fixtures, DoD summary for a story or feature. Parallel subagents.
RVbmad-testarch-test-reviewReview quality of written tests — audit against the knowledge base and best practices, deliver specific improvements.
NRbmad-testarch-nfrNFR Evidence Audit — assess implemented NFR evidence, recommend missing actions.
TRbmad-testarch-traceTrace coverage & Release Gate — map requirements → tests (5 steps), final step yields a decision: PASS / CONCERNS / FAIL / WAIVED.
GATErouting intent NEW v6.10.0Composite release-gate shortcut — walks optional test-review → optional nfr-assess → trace Phase 2 gate, without merging the underlying workflows.
How to Invoke Murat
Type "talk to Murat" or "Test Architect" to activate the agent. Murat will greet by name, display the menu, and wait for a selection.
If the intent is clear (e.g., "hey Murat, design tests for this epic"), Murat skips the menu and dispatches directly to the appropriate skill.
The 🧪 prefix in every response indicates Murat is active.
§
Skills — Teach me testing
04 · TMT
Teach Me Testing TEA
Interactive testing academy — 7 role-adapted sessions from fundamentals to advanced patterns. Adapts depth to learner role (QA / Dev / Lead / VP) and experience level.
run byLearner
  1. 01
    Assess & Profile
    Gather role, experience level, and learning goals. Check for an existing progress file.
    onboard
    what

    Collect role (QA / Dev / Lead / VP), experience level (Beginner / Intermediate / Experienced), learning goals, and optional pain points. Detect existing progress file.

    why

    Depth and examples adapt per role — a QA needs practice focus, a VP needs strategy and metrics context.

    how

    Load or create {user}-tea-progress.yaml; pre-select recommended session path based on experience.

  2. 02
    Select Session
    Choose from the 7-session hub menu; return here between sessions.
    navigate
    what

    Present 7 sessions with status (✅ completed / 🔄 in-progress / ⬜ not-started), scores, and duration. Allow non-linear selection.

    why

    Non-linear navigation lets experienced learners jump directly to advanced content without re-covering basics.

    how

    Select 1–7 to enter a session; X exits and saves progress; all-complete routes automatically to certificate generation.

  3. 03
    Teach
    Deliver role-adapted content for the chosen session using TEA knowledge fragments.
    learn
    what

    Present concepts, examples, and TEA knowledge fragments adapted to learner role. Sessions cover: Quick Start → Core Concepts → Architecture → Test Design → ATDD & Automate → Quality & Trace → Advanced Patterns.

    why

    Just-in-time content loading keeps each session focused and context-window manageable across the full 24k-line TEA knowledge base.

    how

    Load session-content-map for this session; select role-appropriate examples; reference local docs and knowledge fragments; web-browsing as fallback for updated framework docs.

  4. 04
    Quiz & Validate
    Knowledge check after each session — ≥70% to advance.
    validate
    what

    3 targeted questions per session with constructive, role-aware feedback. Up to 3 attempts; score saved to progress file.

    why

    Active recall cements learning and surfaces gaps before the learner moves on to the next session.

    how

    Score ≥70% proceeds; <70% offers focused review of the specific topics that failed. Session 7 (Advanced Patterns) is exploratory — auto-scores 100% on completion.

  5. 05
    Generate Artifacts
    Save session notes, update progress file; issue completion certificate when all sessions are done.
    artifact
    what

    Write session-notes.md from template; update {user}-tea-progress.yaml (status, score, completion %). Generate completion certificate when all 7 sessions reach ≥70%.

    why

    Learners keep structured notes and a persistent progress record that survives across days and conversations.

    how

    Auto-save after each quiz completion. Certificate includes all session scores, skills acquired, and next-step recommendations.

IN
role + experience levellearning goals
OUT
progress YAMLsession notescompletion certificate
Unified Workflow
BMM + WDS + TEA — every phase in one view
BMM + WDS + TEA Unified Workflow Diagram
01
Analysis
Explore the problem space · Validate ideas before planning
Optional
01 · A
Brainstorming
Guided brainstorming session targeting 100+ ideas, automatically shifting perspective every 10 ideas. 3 selectable stances (Facilitator / Creative Partner / Ideate-for-me), a visual technique composer (brain-selector.html, 108 techniques), and an append-only memlog with optional --by user/--by coach authorship attribution.
managed byAnalyst
  1. 01
    Select technique
    Mind-map, SCAMPER, reverse thinking — choose the method that fits the context.
    mind mappingSCAMPER
    what

    AI initializes the session with a 3-question interactive setup then presents a technique selection menu.

    • 1
      AI asks 3 setup questions to define session scope
      • topic — main topic of the brainstorm
      • goal — broad exploration or focused ideation (deep)
      • capture session — whether to save the output document (default Yes)
    • 2
      AI offers 4 technique selection modes
      • Mode 1 — User-selected: user browses library of ~60 techniques and picks
      • Mode 2 — AI-recommended: AI analyzes context, suggests 3-4 best fits with one-line rationale
      • Mode 3 — Random: AI picks randomly (good for breaking patterns)
      • Mode 4 — Progressive flow: moves sequentially from broad → narrow through multiple techniques
    • 3
      User confirms choice before moving to step 02 — no auto-progress
    why
    • ·
      No suitable technique → session becomes unfocused or falls into anchor bias
    • ·
      Each technique opens a different angle of attack on the problem space
    • ·
      SCAMPER is effective when improving existing solutions (Substitute / Combine / Adapt / Modify / Put to other use / Eliminate / Reverse)
    • ·
      Reverse Brainstorming is effective when stuck — ask "how to create the problem?" instead of "how to solve it?"
    • ·
      Six Thinking Hats is effective when balancing multiple perspectives is needed (facts / emotions / risks / benefits / creative / process)
    • ·
      Mind Mapping is effective for broad exploration without a specific scope
    how

    AI loads the technique library (~60 entries) from the internal knowledge base then routes by mode.

    • 1
      User-selected mode: AI displays the full list with category labels (divergent / convergent / lateral / structured)
    • 2
      AI-recommended mode: AI analyzes 2 axes
      • user context (industry, problem type, stage)
      • strengths of each technique (internal matching matrix)
      • returns top 3-4 techniques with a one-line justification
    • 3
      After user picks, AI confirms with a summary: "OK, we'll use [technique] for [goal]. Ready to start?"
  2. 02
    Generate ideas in batches
    Target 100+ ideas, quantity-first principle, no judgment during generation.
    divergent thinkingtimeboxing
    what

    AI opens the generation session, prompts according to the chosen technique, generates ideas in batches of 10-15 ideas/round, targeting 100+ ideas total.

    • 1
      AI prompts the user according to the technique template (e.g., SCAMPER: "What can we Substitute in the current design?")
    • 2
      Each round is timeboxed ~5-10 minutes, AI applies 3 principles
      • defer judgment — no evaluation of any idea during this round
      • build on others — combine existing ideas into new ones
      • wild ideas welcome — outrageous ideas are encouraged
    • 3
      After each round, AI asks for brief feedback ("good direction?") then continues
    • 4
      AI does not self-generate ideas but pulls ideas out of the user via probing questions
    • 5
      Track count continuously, target 100+ ideas before moving to convergence
    why
    • ·
      Per creativity research: the first 20 ideas are usually obvious, ideas 50-100 are the "golden" zone
    • ·
      Early critique kills divergent thinking — weak ideas still have value as stepping stones to good ones
    • ·
      Quantity-first breaks self-censorship — users verbalize minor ideas more easily when they know filtering comes later
    • ·
      Timeboxed rounds maintain energy — open-ended sessions tend to degrade in quality after 30 minutes
    how
    • ·
      AI uses probing questions such as: "What if cost was zero?", "What would Tesla do?", "What's the opposite approach?"
    • ·
      AI uses analogical prompting: "What if this were like X industry?" to break domain anchoring
    • ·
      Each idea is recorded with a 1-line description, no expansion required — speed > polish at this stage
    • ·
      AI counts silently and signals when milestones are reached ("50 ideas reached, keep going")
  3. 03
    Anti-bias pivot
    Every 10 ideas, shift perspective / persona to break cognitive bias.
    cognitive biassix thinking hats
    what

    After every ~10 ideas, AI actively shifts the lens to break patterns and open new idea categories.

    • 1
      AI tracks session metadata: idea count, which technique, which cluster is dominating
    • 2
      When threshold is reached (every 10/20/30 ideas), AI self-pivots using one of 3 methods
      • Persona shift: "Now think as a CFO / kid / hacker / regulator"
      • Hat shift: switch Six Thinking Hat (white = facts → red = emotions → black = risks → green = creative...)
      • Question inversion: "How would we make this feature FAIL?"
    • 3
      AI does not ask permission but declares "pivot now — let's look from X angle" and continues
    • 4
      After a pivot, new-round ideas are often more outlier → flag separately to avoid early merging
    why
    • ·
      Anchoring bias: the brain anchors to the initial framing — pivots break that anchor
    • ·
      Availability heuristic: the next idea tends to be close to the last one → homogeneous clusters
    • ·
      Functional fixedness: only seeing 1 use case per thing — persona shifts open alternative use cases
    • ·
      Active pivots (without waiting for user request) create structured serendipity
    how
    • ·
      AI uses an internal perspective rotation matrix: 6 personas × 3 timeframes × 4 constraint levels
    • ·
      Persona shifts in a 60-minute session: 3-4 different personas, ~15-20 ideas per persona
    • ·
      Standard Six Hats sequence: white (data) → red (feelings) → black (risk) → yellow (benefit) → green (creative) → blue (process)
    • ·
      Question inversion technique from Reverse Brainstorming: list ways to cause the problem, then negate them
  4. 04
    Cluster ideas
    Group ideas by theme, name clusters, remove duplicates.
    affinity mappingclustering
    what

    Convergence mode begins. AI groups ideas into clusters by affinity, names categories, and deduplicates.

    • 1
      AI scans all 100+ ideas and proposes 5-9 initial clusters
    • 2
      Cluster names use "categories everyone understands"
      • e.g., UX improvements, Revenue models, Tech debt, Compliance
      • avoid jargon or team-internal names
    • 3
      User approves or renames each cluster
    • 4
      AI assigns each idea to exactly 1 cluster
    • 5
      Outlier ideas (one-of-a-kind, don't fit any cluster)
      • placed in a separate parking lot
      • never forced into a cluster (preserves originality)
    • 6
      Duplicate ideas merged by content (merge count → signals cluster importance)
    why
    • ·
      100+ spread-out ideas are noise — not actionable, not communicable
    • ·
      Clustering turns noise into signal — reveals themes the session naturally gravitated toward
    • ·
      "Populous" themes are usually the most investment-worthy areas of the problem space (revealed preference)
    • ·
      "Sparse" themes also carry information — gaps may be blind spots or less important areas
    • ·
      Per Miller's Law (7±2): human short-term memory holds ~5-9 items — clusters in this range are easier to retain
    how
    • ·
      AI uses affinity mapping protocol: physical metaphor (sticky notes on a wall) even in digital form
    • ·
      Clustering is conceptual — by purpose/audience/mechanism, not by similar wording
    • ·
      After clustering, count ideas/cluster → bar chart visual so the user can see distribution
    • ·
      Parking lot outlier ideas are often material for a future epic — do not discard
  5. 05
    Prioritize
    Assess impact vs. effort, select top ideas to develop further.
    impact/effort matrixMoSCoW
    what

    AI applies a prioritization framework, scores ideas, and selects the top 5-10 to carry into Phase 2.

    • 1
      AI offers 2 prioritization frameworks
      • Impact/Effort 2×2 matrix: quick wins / big bets / fill-ins / thankless tasks
      • MoSCoW: Must have / Should have / Could have / Won't have
    • 2
      User selects framework (default: Impact/Effort — more visual)
    • 3
      Scoring process
      • Impact (1-5): "if shipped, how much user/business value does this deliver?"
      • Effort (team-weeks): "how long to build and ship?"
      • AI proposes initial scores, user adjusts
    • 4
      Plot ideas on the visual matrix
    • 5
      Output: top 5-10 ideas with short rationale (1-2 lines each)
    • 6
      Quick wins (high impact + low effort) flagged separately for immediate action
    why
    • ·
      A great brainstorm without prioritization = a dead document
    • ·
      Prioritization forces explicit trade-offs and prevents cherry-picking later
    • ·
      Scoring surfaces hidden disagreement: if 2 people score very differently, different assumptions need to be clarified
    • ·
      Quick wins shipped early create momentum + early evidence for the Phase 2 PRD
    • ·
      Big bets need additional research/validation before commitment — flag to avoid hasty decisions
    how
    • ·
      AI uses Impact/Effort 2×2 in Eisenhower box style — 4 quadrants with pre-defined names
    • ·
      Effort dimension: AI estimates using reference class forecasting — compares against similar features with known history
    • ·
      Impact dimension: AI tries to quantify ("saves 2 hours/user/week") rather than using generic ratings ("high")
    • ·
      Avoid vanity metrics: prioritize retention/engagement over page views/downloads
  6. 06
    Save report
    Export brainstorming-report.md, recommend next workflow.
    documentation
    what

    AI generates brainstorming-report.md containing the full session and saves it to _bmad-output/planning-artifacts/.

    • 1
      Report sections in order
      • Setup: topic, goal, technique(s) used, date, attendees
      • Full idea list: numbered, grouped by round
      • Clusters: with count + name + member ideas
      • Prioritization matrix: visual 2×2 or MoSCoW table
      • Top picks: 5-10 ideas with rationale
      • Next steps: recommended workflow to continue
    • 2
      Frontmatter attaches metadata: stepsCompleted, workflowType: brainstorming, date
    • 3
      AI invokes bmad-help with arg = "brainstorming"
    • 4
      bmad-help suggests next workflow based on project state
      • typically: research (validate ideas) or product-brief (document vision)
      • if ideas are already solid: go straight to create-prd
    why
    • ·
      Output artifact = single source of truth for subsequent phases
    • ·
      Without a file, context is lost when changing chat sessions (core BMAD principle: "artifacts drive state")
    • ·
      Teammates cannot re-derive results from a verbal summary alone
    • ·
      Frontmatter metadata enables workflow continuation — re-invoking a workflow knows which steps are already done
    how
    • ·
      Append-only writing: AI writes section by section, does not regenerate the file
    • ·
      File path resolves via config: {planning_artifacts}/brainstorming-{slug}-{date}.md
    • ·
      bmad-help auto-invokes at the final step of every workflow — uniform design pattern
    • ·
      Roadmapping syntax: next phase, next workflow, prerequisite, estimated time
IN
—
OUT
brainstorming-report.md
01 · B
Research Trio
Domain · Market · Technical research to validate assumptions before commitment.
managed byAnalyst
  1. 01
    Define research questions
    Clarify the specific questions to answer (market size? tech feasibility? compliance?).
    problem framing5 whys
    what

    AI turns a vague question into 5-8 sub-questions answerable with evidence.

    • 1
      AI asks user to select one or more research scopes
      • Market — TAM/SAM/SOM, competitors, trends, pricing
      • Domain — vocabulary, expert models, regulations
      • Technical — feasibility, stack options, implementation paths
    • 2
      AI applies 5 Whys to drill from the original question down to the root question
      • From each answer, AI asks 'why?' again
      • Stop at level 5 or when a researchable question is reached
    • 3
      AI converts root question into 5-8 specific, evidence-answerable sub-questions
    • 4
      User validates the sub-question list before AI begins gathering
    why
    • ·
      A vague question leads to wandering research — specificity creates clear coverage checklists
    • ·
      Small sub-questions enable parallel investigation and a clear signal for 'done'
    • ·
      Surfaces scope creep early — one original question can imply 10 sub-areas
    how
    • ·
      Use the problem framing canvas: problem / users affected / current solutions / why insufficient
    • ·
      Each sub-question must be answerable with evidence, not an opinion question
    • ·
      Standard format: "Are there competitors using on-device LLM for document Q&A, and what is their pricing?"
  2. 02
    Gather sources
    Web search, docs, interviews, industry reports — record sources completely.
    desk researchsource triangulation
    what

    AI gathers evidence from 3+ independent source types, attaching a full citation to each finding.

    • 1
      AI searches in parallel across 4 source types
      • Web search — news, blogs, industry reports
      • Official docs — specs, standards (ISO/RFCs), regulatory texts
      • Expert content — podcasts, talks, academic papers
      • Interview notes — if the user has existing user research
    • 2
      Each finding is assigned a quality grade
      • Primary — regulator docs, SEC filings, official statements
      • Secondary — Reuters/Bloomberg, peer-reviewed research
      • Tertiary — Wikipedia, aggregators, blog opinions
    • 3
      Each citation includes: URL + publish date + author name + verbatim quote or clear paraphrase
    why
    • ·
      Triangulation — a fact appearing in 3 independent sources = evidence; 1 source = assumption
    • ·
      Mandatory citations prevent AI from hallucinating stats — must attach a specific URL
    • ·
      Quality grading helps downstream synthesis weight evidence appropriately
    how
    • ·
      Desk research protocol: primary first, then secondary, tertiary only for cross-checking
    • ·
      Multi-query strategy: rephrase each sub-question 2-3 different ways to cover the keyword space
    • ·
      Recent-first filter: prioritize sources from the past 12 months for fast-changing topics
  3. 03
    Analyze data
    Cross-reference, identify patterns, contradictions, and gaps in the information.
    pattern recognitionSWOT
    what

    AI cross-references data and looks for 3 types of signal: patterns, contradictions, gaps.

    • 1
      AI builds a finding matrix
      • Rows — sub-questions from step 01
      • Columns — sources gathered
      • Cells — quote/data from that source for that question; empty cell = gap
    • 2
      AI scans the matrix for 3 signal types
      • Patterns — ≥3 sources agree on X → high-confidence finding
      • Contradictions — source A says X, source B says ~X → highlight and investigate
      • Gaps — empty cells in the matrix → unanswered questions
    • 3
      AI does not skip contradictions or gaps — both are predictive of risk
    why
    • ·
      Pattern recognition is easy to overweight — contradictions and gaps are often overlooked but carry high value
    • ·
      Contradictions are typically contested ground worth digging into
    • ·
      Gaps are often unknown unknowns — predictive of risks in Phases 3-4
    how
    • ·
      Matrix rendered as a table (sub-questions × sources); empty cells naturally surface as gap indicators
    • ·
      Pattern threshold: ≥3 independent sources = pattern; 2 = weak signal; 1 = assumption
    • ·
      Contradictions: AI presents both sides with evidence weight, does not pick a winner
  4. 04
    Synthesize findings
    Extract key insights, quantify where possible, and highlight risks.
    synthesisinsight distillation
    what

    AI synthesizes into 3-5 actionable insights, quantifying when possible.

    • 1
      Each insight has the format
      • Claim — 1 line, specific, quantified when data is available
      • Evidence pointer — cite which source numbers in the matrix
      • Confidence level — high / medium / low
    • 2
      AI lists 3-5 risk flags cross-category
      • Market — timing, demand, competition
      • Technical — feasibility, scalability, vendor
      • Regulatory — compliance, data, jurisdiction
    • 3
      Output: 1 document readable in 5 minutes by a PM/founder
    why
    • ·
      Synthesis ≠ summary — synthesis creates new understanding by combining ('because X and Y, therefore Z')
    • ·
      Quantifying forces precision — 'growing fast' could be 5% or 50%, a critical difference
    • ·
      Confidence level helps PM weight insights when they conflict with gut feeling
    how
    • ·
      Pyramid Principle (Minto): conclusion-first, supporting evidence below
    • ·
      Confidence calibration: high = 3+ primary sources; medium = 2 sources; low = 1 + inference
    • ·
      Each insight tagged with evidence count (X sources) for quick reviewer judgment
  5. 05
    Recommendations
    Provide specific recommendations based on collected evidence.
    evidence-based
    what

    AI provides evidence-based recommendations — each with action, rationale, owner, metric.

    • 1
      Each recommendation has the format
      • Given — evidence from insights
      • Recommend — specific action
      • Because — reasoning linking evidence → action
      • Owner — PM / Tech Lead / Founder
      • Success metric — how to know it worked
      • Risk if skipped — consequence of inaction
    • 2
      AI avoids opinion-only — each recommendation must trace back to evidence from steps 03/04
    • 3
      If evidence is weak, AI flags low-confidence recommendation instead of pretending it's solid
    why
    • ·
      Research without recommendations = trivia — PMs need decisions, not facts
    • ·
      Given/Recommend/Because format makes recommendations defensible in meetings with stakeholders
    • ·
      Risk if skipped creates a 'cost of inaction' — helps prioritize against other proposals
    how
    • ·
      Argumentation theory: claim + grounds (evidence) + warrant (reasoning)
    • ·
      Ranking: impact × confidence × urgency
    • ·
      AI offers 2-3 alternative paths when there is genuine uncertainty instead of forcing 1 answer
  6. 06
    Save research report
    Export research-findings.md with full citations and next actions.
    documentationcitations
    what

    AI generates research-findings.md with full citations in APA/Chicago format.

    • 1
      Report structure
      • Executive summary — 1 paragraph
      • Methodology — scope, sources, date range
      • Sub-questions + evidence — 1 section/question
      • Insights — 3-5 with confidence levels
      • Recommendations — ranked with rationale
      • Full bibliography — alphabetical, complete citations
    • 2
      Frontmatter: dates, methodology, scope, workflowType: research
    • 3
      AI invokes bmad-help → suggest next workflow based on strength of findings
    why
    • ·
      Citations ensure reproducibility — PM can trace every claim back to its source
    • ·
      Methodology section enables scope critique — reviewer can see if key sources are missing
    • ·
      Artifacts drive state — without a file, context is lost when switching chats
    how
    • ·
      Format: APA (Author, Year, Title, URL) or Chicago (with footnotes)
    • ·
      Append-only writing — AI writes each section, does not regenerate the file
    • ·
      End of file includes a 'Last updated' timestamp to track freshness
IN
brainstorming-report.md
OUT
research-findings.md
01 · C
Product Brief
Outcome-driven facilitator (v6.10.0) — 3 intents: Create / Update / Validate. Fast path (batch gaps → draft) + Coaching path (section by section). 3-file output: brief.md + addendum.md + .decision-log.md.
managed byAnalyst
  1. 01
    Init & context
    Load prior artifacts (brainstorming, research), define the brief scope.
    context setting
    what

    AI loads prior artifacts and scopes the brief for the target audience.

    • 1
      AI checks the workspace for Phase 1 artifacts
      • brainstorming-report.md
      • research-findings.md
      • prfaq-{project}.md
    • 2
      If found: AI extracts relevant content — no re-derivation
    • 3
      AI asks 3 scoping questions
      • Audience — founder / engineering / mixed / stakeholder external
      • Decision driven — what decision does this brief inform?
      • Level of detail — 1-2 pages (preferred) or deeper
    • 4
      AI propose outline → user adjust → commit
    why
    • ·
      Brief without scope → becomes a PRD lite — long, rambling
    • ·
      A proper brief must be 1-2 pages, stakeholders read it in 5 minutes for a go/no-go decision
    • ·
      Extracting (not re-deriving) saves time and maintains consistency with earlier artifacts
    how
    • ·
      Default audience: mixed (founder + early team) if user does not specify
    • ·
      Artifact extraction: AI scans frontmatter + headings, lists relevant sections for user confirmation
    • ·
      Outline proposed as bullet list for user to adjust before writing
  2. 02
    Vision & mission
    State the long-term vision and product mission concisely.
    vision crafting
    what

    AI co-drafts Vision (5-10 years) and Mission (1 sentence, what-for-whom-how).

    • 1
      Vision drafting
      • AI ask: 'What does world look like in 5-10 years if you succeed?'
      • AI propose 2-3 draft options
      • User selects + iterates until no more words are needed
    • 2
      Mission drafting with the Geoffrey Moore template
      • 'For [user], who [pain], our [product] is [category] that [value]. Unlike [alternative], we [differentiator].'
      • Compress down to the final 1-2 sentences
    • 3
      Validate: vision is distant from tactics; mission is close to tactics
    why
    • ·
      Vision alone = a dream with no daily direction
    • ·
      Mission alone = labor without purpose, losing long-term direction
    • ·
      The two are complementary and serve as an anchor whenever feature creep occurs
    how
    • ·
      Opposite test: if the opposite is not plausible → vision is too bland, redo
    • ·
      Read aloud test: if it sounds generic / applies to any startup → re-draft
    • ·
      Vision/Mission must be finalized before step 03 — all future decisions reference back to them
  3. 03
    Target users
    Sketch 1-3 primary personas with specific needs and pain points.
    personasjobs-to-be-done
    what

    AI sketches 1-3 personas: demographics, JTBD, pain, current alternative, success criteria.

    • 1
      Each persona format
      • Demographics — age, role, location, tech savviness
      • JTBD — 'When [situation], I want to [motivation], so I can [outcome]'
      • Primary pain — 1-2 specific lines
      • Current alternative — how they solve it today (including 'do nothing')
      • Success criteria — measurable outcome if pain resolved
    • 2
      AI does not merge all users into 1 'average user'
    • 3
      Validate with user
      • 'Do you know a real person who matches this persona?'
      • If not → persona may be fiction, redo
    why
    • ·
      Average user is fiction — real users have different jobs/pain
    • ·
      Products trying to serve everyone usually serve no one
    • ·
      JTBD framework focus on motivation rather than demographic stereotype
    how
    • ·
      AI uses Clayton Christensen's JTBD framework
    • ·
      AI uses the empathy map as a supplementary tool: says / thinks / does / feels
    • ·
      Also define the anti-persona — who we explicitly do not serve
  4. 04
    Problem statement
    Define the problem being solved, why it matters and why it's urgent.
    problem framing
    what

    AI writes a Problem Statement specific enough to be falsifiable.

    • 1
      4 components of the problem statement
      • What — specific problem, not a solution disguised as a problem
      • Who — who is affected, link to personas
      • Why important — frequency × severity × visibility
      • Why now — timing: regulation, tech enablement, market shift
    • 2
      AI avoids solution-in-the-problem (e.g., 'we need a chatbot' = solution, not a problem)
    • 3
      Statement must be falsifiable: can be refuted by evidence
    why
    • ·
      Vague problem → vague solution → shipping features that don't address the core issue
    • ·
      The problem statement is a contract between the team and reality
    • ·
      Being specific enables measuring real success, not just 'feature shipped'
    how
    • ·
      PAIN framework: frequency × severity × visibility
    • ·
      Falsifiability test: 'What evidence would prove this problem doesn't exist?'
    • ·
      AI uses 5 Whys to dig past surface complaints to the root problem
  5. 05
    Value proposition
    Unique value delivered to users, differentiated from existing solutions.
    value prop canvas
    what

    AI defines the Value Proposition: unique value vs. alternatives, mapping pains → gains.

    • 1
      AI uses the Value Proposition Canvas
      • Customer pains ← linked to problem statement
      • Customer gains — positive outcomes desired
      • Pain relievers — how product addresses each pain
      • Gain creators — how product creates each gain
    • 2
      AI maps value vs. 3 alternatives: top competitors + current workaround + 'do nothing'
    • 3
      Articulate differentiator: 1-2 'only we have because [moat]'
    why
    • ·
      Most products fail not because they are bad, but because they are not different enough for users to switch
    • ·
      'Do nothing' is always the strongest competitor — switching cost > pain = no adoption
    • ·
      Articulating the value prop early reveals: if you cannot articulate it, it may not exist
    how
    • ·
      Value Proposition Canvas by Strategyzer/Osterwalder
    • ·
      Differentiation matrix: rows = features, cols = us + top 3 competitors
    • ·
      AI flags features that all competitors have = table stakes, not differentiators
  6. 06
    Success metrics
    KPIs and OKRs to measure success — short-term and long-term.
    OKRnorth star metric
    what

    AI defines Success Metrics: 1-2 north star + 3-5 supporting KPIs.

    • 1
      North Star Metric — 1-2 metrics that, if improved, improve the entire business
      • Airbnb = nights booked; Spotify = time listening; Slack = messages sent
    • 2
      Supporting KPIs per the HEART framework
      • Happiness — NPS, satisfaction
      • Engagement — DAU/MAU, session frequency
      • Adoption — signup conversion, time-to-first-value
      • Retention — D7/D30, cohort curves
      • Task success — completion rate, error rate
    • 3
      Separate timeframes
      • Short-term (3-6 months) — leading indicators
      • Long-term (1+ year) — lagging indicators
    why
    • ·
      Metrics are a contract between the team and the board
    • ·
      Vanity metrics (page views, downloads) are easy to grow but do not correlate with business value
    • ·
      The North Star aligns the entire team around a single number
    how
    • ·
      OKR template + HEART framework by Google
    • ·
      Each metric: baseline (current) + target (12 weeks) + measurement source
    • ·
      AI flags a warning if a metric has no baseline/target → 'vague metric'
    • ·
      Layer North Star + 2-3 leading indicators — avoid putting all expectations on a single number
  7. 07
    Constraints
    Budget, timeline, resources, technical and regulatory constraints.
    constraint analysis
    what

    AI lists Constraints: budget, timeline, team, technical, regulatory — categorized as MUST/SHOULD/MAY.

    • 1
      AI asks the user about 5 types of constraints
      • Budget — dollar amount, runway
      • Timeline — launch date, milestones
      • Team — size, skills, hiring plan
      • Technical — must integrate with X, must use Y stack
      • Regulatory — HIPAA / GDPR / SOC2 / PCI / industry-specific
    • 2
      Each constraint categorized into 3 tiers
      • MUST — deal-breaker if violated
      • SHOULD — preferred
      • MAY — nice-to-have
    • 3
      AI flags conflicts between 2 MUST constraints — must be resolved before committing
    why
    • ·
      Constraints define solution space — not listing them = assuming they don't exist = surprise mid-sprint
    • ·
      A constraint can become a differentiator (e.g., 'on-device' due to privacy constraint becomes a USP)
    • ·
      MUST constraint conflicts must be resolved before committing — not mid-development
    how
    • ·
      Constraint analysis matrix: type × tier × source
    • ·
      AI asks regulatory questions per industry: healthcare → HIPAA, finance → SOC2/PCI, EU → GDPR
    • ·
      Brand constraints are often forgotten: voice, visual identity, tone restrictions
  8. 08
    Next steps
    Save product-brief.md and recommend moving to Phase 2 (PRD).
    roadmapping
    what

    AI saves product-brief.md and clearly hands off to Phase 2.

    • 1
      Final document structure
      • Executive summary — 1 paragraph (written last, placed first)
      • Vision + Mission
      • Target users — 1-3 personas
      • Problem statement
      • Value proposition + differentiation
      • Success metrics — north star + KPIs
      • Constraints — tier categorized
    • 2
      Frontmatter: workflowType: product-brief, stepsCompleted complete
    • 3
      AI invokes bmad-help → recommends next: create-prd or prfaq
    why
    • ·
      The brief is the input for the PRD — the PRD source-extracts the brief automatically, no re-explaining
    • ·
      Executive summary written last but placed first per the Pyramid Principle
    • ·
      1-2 page format enforces tight writing — quality over quantity
    how
    • ·
      Append-only writing: write each section, do not regenerate
    • ·
      File path: {planning_artifacts}/product-brief.md
    • ·
      bmad-help auto-invoked at the final step — uniform design pattern for all workflows
IN
brainstorming-report.mdresearch-findings.md
OUT
product-brief.md
01 · D
PRFAQ Challenge
Working Backwards per Amazon — Press Release, External & Internal FAQ, CX, Validation.
managed byAnalyst
  1. 01
    Press Release
    Write a hypothetical press release for launch day — headline, subhead, customer quotes.
    working backwardsnarrative writing
    what

    AI writes a hypothetical press release for launch day in AP/Reuters style.

    • 1
      PR structure (~1 page)
      • Headline — 8-12 words, attention-grabbing
      • Subhead — 1 sentence expand on headline
      • Lead paragraph — who/what/when/where/why
      • Body — 2-3 paragraphs detail
      • Customer quotes — 2-3 concrete, specific quotes (e.g., 'saved me 4h/week')
      • Leadership quote — articulate the theory of change
      • Call-to-action — where to sign up / learn more
    • 2
      AI writes as a freelance journalist — if it feels boring, the product is not exciting enough yet
    why
    • ·
      Amazon's Working Backwards — Jeff Bezos's methodology
    • ·
      If you cannot write a compelling PR, the product is not ready to build
    • ·
      PR forces articulating the value proposition in user-facing language, not engineer-facing
    • ·
      Customer quotes force specific outcomes — vague benefits = vague product
    how
    • ·
      Inverted pyramid: most important info top, details bottom
    • ·
      Boredom test: read aloud — if it feels cringe, redo
    • ·
      Customer quotes format: '[Product] saved me [specific outcome]' — not 'I love it'
  2. 02
    External FAQ
    FAQ for customers: how to use, pricing, benefits, comparison with other products.
    customer empathy
    what

    AI generates External FAQ (8-12 Q/A) from the perspective of a skeptical customer.

    • 1
      FAQ covers the most uncomfortable topics
      • Pricing — 'Why this expensive?', 'Free tier?'
      • Comparison — 'How different from [competitor]?'
      • Risk — 'Data privacy?', 'Vendor lock-in?'
      • Migration — 'How to migrate from current solution?'
      • Support — 'What if it breaks?', 'Refund policy?'
    • 2
      The hardest questions at the top of the FAQ — not buried at the bottom
    • 3
      Each answer 2-4 lines: direct + concrete, no marketing fluff
    why
    • ·
      FAQs expose uncomfortable questions the team typically avoids
    • ·
      Answering before launch = ready to defend the product; unable to answer = not yet solid
    • ·
      Hardest questions on top — if the team is not sure = must resolve before committing
    how
    • ·
      Devil's advocate prompting: AI generate FAQ as if hostile to product
    • ·
      Answer format: direct + evidence + escape hatch if applicable
    • ·
      Avoid marketing fluff — concrete numbers/comparisons rather than adjectives
  3. 03
    Internal FAQ
    Internal FAQ: risks, assumptions, costs, timeline, technical trade-offs.
    pre-mortemrisk register
    what

    AI generates the Internal FAQ for the team/board: risks, costs, assumptions, failure modes.

    • 1
      Internal FAQ covers
      • Top risks — market, tech, team, regulatory, financial
      • Cost breakdown — dev cost, COGS, marketing CAC
      • Timeline assumptions — '6 months — what if it takes 12?'
      • Technical trade-offs — chosen path vs alternatives
      • 'What if we're wrong about X' — pre-mortem questions
    • 2
      AI uses the pre-mortem protocol
      • Assume product fails in 12 months — tell the story of WHY
      • Top 5 failure reasons → converted into a risk register
    why
    • ·
      Internal FAQ is the pre-mortem mechanism
    • ·
      Per Kahneman: pre-mortem improves forecast accuracy by 30%
    • ·
      Surface hidden assumptions before committing — after committing it costs 10-100x to fix
    how
    • ·
      Each risk: probability × impact + early warning signal + mitigation
    • ·
      AI prompts hardest hypotheticals: 'What if Apple does this in 6 months?'
    • ·
      Risk categories: Market / Technical / Team / Regulatory / Financial
  4. 04
    Customer Experience
    Narrate the detailed end-to-end experience from the user's perspective (moments of truth).
    service blueprintuser journey
    what

    AI narrates the Customer Experience end-to-end: awareness → habit → renewal.

    • 1
      5 stages with target emotion
      • Awareness — how user hears about product
      • Consideration — what convinces them to try
      • Onboarding — first 5 minutes after signup
      • First use — aha moment / value realization
      • Habitual use — daily/weekly engagement loop
    • 2
      AI identifies moments of truth at each transition
      • Critical decision points where user can drop off
      • Typically at: pricing page, first-use friction, error recovery
    why
    • ·
      A static features list says 'what exists'; CX narrative says 'how it feels'
    • ·
      UX problems mostly occur at transitions between stages, not in the stages themselves
    • ·
      Peak-end rule (Kahneman): users remember peak emotion + end emotion → design accordingly
    how
    • ·
      Service blueprint: above-the-line (user-visible) + below-the-line (backend)
    • ·
      User journey mapping with emotion curve
    • ·
      Deliberately create 1 peak emotion + strong end emotion
  5. 05
    Validation
    Review everything, stress-test with hard questions, finalize the prfaq file.
    stress testingpeer review
    what

    AI stress-tests the entire PR/FAQ from 5 adversarial perspectives, patches gaps, finalizes.

    • 1
      AI plays adversarial reviewer from 5 angles
      • Investor — 'TAM? Unit economics? Moat?'
      • Customer — 'Why would I switch? Risk?'
      • Engineer — 'Feasible at scale?'
      • Competitor — 'How quickly can we copy?'
      • Regulator — 'What compliance is needed?'
    • 2
      Each question → grade the answer
      • Strong — keep as-is
      • Weak — patch/strengthen the related section
      • No answer — flag as gap, document for later
    • 3
      Finalize prfaq-{project}.md → invoke bmad-help → suggest create-prd
    why
    • ·
      Stress test = attempting to break the product on paper before breaking it in reality
    • ·
      If the PR/FAQ withstands 10 adversarial questions, the concept has resilience
    • ·
      Gap documentation → research backlog for Phase 2
    how
    • ·
      AI cycles through 5 personas, 2-3 hard questions per persona
    • ·
      Iterate until all answers are ≥ 'strong' or gaps are explicitly documented
    • ·
      Adversarial review pattern reused in Phase 4 Code Review (04·F)
IN
product-brief.md
OUT
prfaq-{project}.md
WDS · 0
Alignment & Signoff WDS
Secure formal stakeholder alignment before any design or development begins — the alignment pitch + optional contract. WDS Phase 0 · optional.
managed bySaga 📚 · WDS Analyst
IN
idea + stakeholders
OUT
A-Product-Brief/pitch.md
WDS · 1
Project Brief WDS
The North Star. Establishes the complete strategic foundation every downstream WDS phase reads. Workshop / Suggest / Dream modes; 36-step flow. WDS Phase 1.
managed bySaga 📚 · WDS Analyst
IN
pitch.md / PRFAQ / PRD
OUT
A-Product-Brief/project-brief.md
WDS · 2
Trigger Mapping WDS
Connects business goals to user psychology via Effect Mapping — positive & negative driving forces → feature impact matrix + personas. WDS Phase 2. Hands the Trigger Map to Freya.
managed bySaga 📚 · WDS Analyst
IN
project-brief.md
OUT
B-Trigger-Map/trigger-map.mdpersonas/
02
Planning
Define what to build and for whom · Requirements & user experience
Required
02 · A
PRD
Single intent-detecting bmad-prd skill — Create / Validate / Update. Essential Spine (§0–§9 always) + Adapt-In Menu (7 groups by concern). Adds .memlog.md audit trail and a Reviewer Gate with HTML report. The former create/edit/validate PRD skills are now deprecated forwarding shims.
managed byPM
  1. 01
    Discovery phase
    Source scan → Brain dump → Artifact scan → Stakes calibration (L0–L4, Hobby → Regulated) → Working mode (Fast / Coaching).
    brain-dumpstakes-calibration
    what

    5 discovery stages before drafting begins.

    • 1
      Source scan — auto-scan planning_artifacts/; surface paths, user confirms before reading
    • 2
      Brain dump — user provides all context in one pass; subagent web research spawns in parallel
    • 3
      Artifact scan — parallel subagents extract decision signals; parent receives relevance-filtered digest
    • 4
      Stakes calibration — Hobby/solo · Internal tool · Consumer launch · Regulated → sets depth for every §, classifies L0–L4
    • 5
      Working mode — Fast Path (batch gaps → draft + [ASSUMPTION] tags) or Coaching Path (section by section); switchable at any point
    why
    • ·
      Stakes calibration sets NFR thresholds, compliance flags, accessibility floor — determines depth of every §
    • ·
      Single brain-dump pass prevents scattered interruptions throughout the session
    • ·
      Early working-mode selection lets PM control pace vs. quality trade-off
    how
    • ·
      Invoke: /bmad-agent-pm → CP (Create), VP (Validate), EP (Edit)
    • ·
      Stakes probe: 1 question, 4 options — no spam; single calibration determines the whole session depth
  2. 02
    §1 · Vision Always
    2–3 paragraphs: what the product is, who it serves, why it matters. Blue Ocean lens for launch-level. Pyramid Principle structure.
    positioningblue-ocean
    what

    2–3 paragraphs compelling enough to stand alone. 3–5 year direction for launch-level; tighter for internal tools. Blue Ocean lens: Eliminate · Reduce · Raise · Create vs. current alternatives.

    why

    Engineering needs 'why now' and 'why us' — not just 'what'. A decision-maker must understand the go/no-go rationale from §1 alone.

    how
    • ·
      Ask: 'Where do you want this product to be in 3 years?' — connect to company strategy
    • ·
      Pyramid Principle: lead with conclusion → supporting argument → detail
  3. 03
    §2 · Target User Always
    §2.1 JTBDs (functional/emotional/social/contextual) · §2.2 Non-Users v1 · §2.3 Key User Journeys (UJ-1…N, named personas, 3–5 beats each).
    JTBDuser-journeys
    what
    • 1
      §2.1 JTBDs — 4 types: functional · emotional · social · contextual. Format: 'When [situation], I want [motivation] so I can [outcome].'
    • 2
      §2.2 Non-Users (v1) — who this is explicitly NOT for. Omitting this invites scope creep.
    • 3
      §2.3 Key User Journeys — named-persona narratives, numbered UJ-1…N. Each UJ: persona + context → entry → path (3–5 beats) → climax (value delivered) → resolution. FRs reference UJ-IDs inline.
    why

    Named personas ('Linh, bookkeeper at XYZ') are mandatory — 'power user' is not actionable. Non-Users establish explicit out-of-scope audience.

    how
    • ·
      Interview ≥1 real user per persona before the session
    • ·
      Narrate a real session → structure into UJ-N form
  4. 04
    §3 · Glossary Always · New v6.9
    Every domain noun defined once. No synonyms elsewhere. Agent builds automatically while drafting. Downstream workflows use Glossary terms as anchors.
    glossarynew-v6.9
    what

    One definition per term — if §4 introduces a new domain noun, it gets added to the Glossary in the same editing pass. New in v6.9: downstream workflows extract source content using Glossary terms as anchors.

    why

    Consistent vocabulary prevents 'is a client the same as an account?' across UX, architecture, and dev teams.

    how

    Agent builds automatically — no dedicated session. Drop when entire readership shares vocabulary. Keep for mixed audiences or regulated domains with precise legal terms.

  5. 05
    §4 · Features Always
    FRs numbered globally FR-1…N. Format: [Actor] can [capability] [conditions]. Realizes UJ-X. Consequences: testable. [ASSUMPTION] inline. Tech choices → addendum.md.
    functional-reqtraceability
    what

    Features grouped by function; each group opens with a behavioral narrative. FRs nested under feature, numbered globally. [ASSUMPTION: ...] tags embedded inline when agent inferred without confirmation.

    why
    • ·
      Numbered FRs = backbone of downstream traceability — architecture, stories, QA all reference FR-IDs
    • ·
      Testable Consequences replace acceptance criteria — each FR must pass without ambiguous interpretation
    how
    • ·
      Ask: 'If we could only ship 5 features, which 5 would you refuse to launch without?'
    • ·
      Write 'what', not 'how' — tech choices belong in addendum.md
  6. 06
    §5 · Non-Goals Always
    What the product is NOT and will NOT do in v1. Three types: in-scope / out-of-scope / future scope. Sign-off required before drafting FRs.
    scope-boundary
    what

    Bulleted. [NON-GOAL for MVP] callouts inline within §4 cover deferred items within features; this section captures the broader 'we are not building X / we are not becoming Y' statements.

    why

    Prevents 'let me add this nearby thing' scope creep at every level. Unwritten out-of-scope items become future misunderstandings.

    how
    • ·
      Ask: 'What features have we discussed that are definitely NOT happening in this release?'
    • ·
      Do not proceed to drafting FRs without explicit sign-off on scope boundaries
  7. 07
    §6 · MVP Scope Always
    In Scope (bulleted) + Out of Scope (each item: reason + defer target: v2, v3, research track). [NOTE FOR PM] for emotionally load-bearing items.
    mvpphasing
    what
    • ·
      In Scope — bulleted, crisp
    • ·
      Out of Scope for MVP — each item has a 1-line reason + explicit defer target (v2, v3, research track)
    why

    MVP ≠ Phase 1: MVP = minimum usable + shippable ('would real users pay for / rely on this?'). A Phase 1 containing 80% of features is not a phase — it's the whole product with a label. False phasing creates false deadline pressure.

    how
    • ·
      Ask: 'If we had to ship in half the time, what would we cut?'
    • ·
      Dependency blockers and half-time cuts reveal the true MVP
  8. 08
    §7 · Success Metrics Always · Counter-metrics new v6.9
    Primary (core thesis) + Secondary (supporting) + Counter-metrics (what NOT to optimize — prevents gaming). Each SM cross-references FR(s) it validates. Baseline required.
    OKRcounter-metrics
    what

    3 tiers: Primary (validates core thesis) · Secondary (supporting signals) · Counter-metrics (prevents gaming the primary — new in v6.9, load-bearing). Length scales with stakes.

    why
    • ·
      Counter-metrics tell the architect what NOT to optimize. Example: maximizing completion rate could tank accuracy.
    • ·
      Primary must be an outcome metric (user value), not an output metric (features shipped)
    how
    • ·
      Every metric needs a baseline: 'from 45 minutes to under 5 minutes'. 'Reduce time significantly' is not a metric.
    • ·
      Ask: 'What number would you look at 6 months after launch to decide if this was worth building?'
  9. 09
    §8 · Open Questions Always
    Numbered list of unresolved issues. Become future tickets or follow-up research — not silent gaps in the document.
    open-questions
    what

    Numbered list. Everything still unknown, unresolved in the session. Agent flags every ambiguity here rather than silently assuming.

    why

    Explicit open questions → tracked tickets. Silent gaps → late-stage surprises.

    how

    Collected during drafting; surfaced before Finalize for PM to triage.

  10. 10
    §9 · Assumptions Index Always
    Every [ASSUMPTION] tag embedded inline surfaced and indexed here. Every index entry must match an inline tag — mismatch flagged as Critical in Reviewer Gate.
    assumptions
    what

    Index of every [ASSUMPTION: ...] tag from throughout the PRD. Agent collects during drafting; PM confirms explicitly before Finalize.

    why

    Untested assumptions = Phase 3-4 surprises. The index makes them visible and actionable.

    how

    Reviewer Gate cross-checks index vs. inline — mismatch is flagged as a Critical finding.

  11. 11
    Adapt-In Menu Optional
    7 optional groups, identified by Concern Scan: (1) Cross-cutting NFRs · (2) Consumer/branded · (3) Enterprise · (4) Regulated domains (13 industries) · (5) Developer products · (6) Embedded/hardware · (7) Small-scope stories.
    adapt-inconcern-scanoptional
    what

    Concern Scan runs automatically as the agent reads Discovery inputs — identifies concerns this product carries and activates the corresponding clusters.

    • 1
      Group 1 — Cross-cutting: Cross-cutting NFRs, Constraints & Guardrails, Why Now
    • 2
      Group 2 — Consumer/branded: Aesthetic & Tone, Information Architecture, Monetization, Platform
    • 3
      Group 3 — Enterprise: Stakeholders & Approvals, Risk & Mitigations, ROI, Operational Requirements, Integration & Dependencies, Rollout, Data Governance, Audit Trail
    • 4
      Group 4 — Regulated domains: Compliance & Regulatory (13-industry taxonomy)
    • 5
      Group 5 — Developer products: API Contracts, Versioning & Deprecation, Performance Budgets, Language/Runtime Targets
    • 6
      Group 6 — Embedded/hardware: Hardware Constraints, Deployment & Update, Environmental & Reliability
    • 7
      Group 7 — Small-scope: Stories (1–2 story max; light PRD + actionable stories in one doc)
    why

    Essential Spine §0–§9 = mandatory. Adapt-In = optional sections added only when the product actually carries that concern — no bloating PRD with irrelevant sections.

    how
    • ·
      Concern Scan is automatic — agent flags applicable clusters from Discovery inputs
    • ·
      When a product carries a concern no cluster names → agent invents the section
  12. 12
    Finalize — 8 steps
    Decision log audit → Input reconciliation → Reviewer pass (Critical/High first) → Triage open items → Polish (structure before prose) → External handoffs → Close → on_complete hook.
    reviewhandoff
    what
    • 1
      Decision log audit — walk .decision-log.md; confirm each entry captured or explicitly set aside
    • 2
      Input reconciliation — subagent per source checks gaps, especially qualitative ideas the FR structure silently drops
    • 3
      Reviewer pass — rubric walker + adversarial reviewer in parallel; findings tiered Critical/High first
    • 4
      Triage open items — [ASSUMPTION] tags, [NOTE FOR PM]; phase-blockers resolved first
    • 5
      Polish — bmad-editorial-review-structure then bmad-editorial-review-prose; structural before prose
    • 6
      External handoffs — Confluence, Notion, Jira...; surface returned URLs/IDs
    • 7
      Close — set status: final in frontmatter; log to .decision-log.md
    • 8
      on_complete hook — run if configured in customize.toml
    why
    • ·
      Decision log prevents re-debate — 6 months later nobody remembers why option B was rejected
    • ·
      Polish structural before prose — do not polish text that will soon be cut
    how
    • ·
      Output files: PRD.md + addendum.md + .memlog.md (append-only audit trail)
    • ·
      Common next steps: bmad-ux · bmad-architecture · bmad-create-epics-and-stories
IN
product-brief.mdprfaq-{project}.md
OUT
PRD.md
02 · B
UX Design
Runs after PRD, before Architecture: user journey, IA, wireframes, interaction, accessibility.
managed byUX Expert
  1. 01
    Load PRD & context
    Read PRD, understand users, FRs and NFRs relevant to UX.
    requirements analysis
    what

    AI loads PRD and extracts UX-relevant FRs + NFRs before starting design.

    • 1
      AI loads prd.md, extracts personas + FRs with UI surface
    • 2
      Identify NFRs that constrain UX
      • Performance — p95 <200ms rules out heavy animations
      • Accessibility — WCAG 2.1 AA requires focus states, contrast ratios
      • Responsive — mobile-first if mobile traffic >50%
    • 3
      Output: FR → UX implication table (e.g.: FR-03 'user uploads file' → drag-drop + progress + error handling)
    why
    • ·
      UX is not separate from requirements — NFR constraints rule out certain design patterns
    • ·
      Starting with requirements load prevents sketching wireframes that need to be revised
    • ·
      Performance NFR in particular: 200ms threshold rules out long skeleton screens
    how
    • ·
      Requirements traceability: FR → UX needs mapping table
    • ·
      AI highlights NFRs that constrain design (mark with ⚠️ in notes)
    • ·
      Personas extracted → primary persona drives primary flow
  2. 02
    User journey mapping
    Map the journey for each persona — touchpoints, emotions, pain points.
    journey mappingJTBD
    what

    AI maps user journey for each persona: touchpoints, emotions, pain points, opportunities.

    • 1
      Journey map 5 stages
      • Awareness — where users hear about product
      • Consideration — evaluation, comparison
      • Onboarding — signup → first meaningful action
      • Regular use — habitual engagement pattern
      • Advocacy — retention, referral, renewal
    • 2
      Each stage: touchpoints + emotion (frustrated/curious/delighted) + pain + opportunity
    • 3
      Flag moments of truth — transitions where users most likely to drop off
    why
    • ·
      Static screens hide flow — journey reveals transitions where UX problems live
    • ·
      Most dropoff happens at transitions, not at the screens themselves
    • ·
      Peak-end rule: design peak delight moment + strong end moment
    how
    • ·
      Journey mapping methodology with an emotion curve
    • ·
      Transitions tagged with 'danger level' — high = prioritize design attention
    • ·
      One journey per persona — don't merge (different people, different paths)
  3. 03
    Information architecture
    Information structure, navigation, sitemap, content hierarchy.
    card sortingsitemap
    what

    AI defines Information Architecture: sitemap, navigation hierarchy, content groupings.

    • 1
      Output of this step
      • Sitemap — all screens/pages and relationships
      • Navigation structure — top-nav / sidebar / breadcrumb decision
      • Content taxonomy — categories, labels, search terms
    • 2
      AI uses card sorting to discover user's mental model
    • 3
      Validate: can the user find feature X in ≤3 clicks?
    why
    • ·
      Bad IA = user gets lost even when the product has enough features
    • ·
      Research: if users can't find it in 3 clicks, they assume the feature doesn't exist
    • ·
      Defining IA upfront = no costly rework after mockups are polished
    how
    • ·
      Card sorting: open (user groups) or closed (user fits into pre-set)
    • ·
      Tree testing: validate IA with real tasks
    • ·
      Output: sitemap diagram + nav structure decision rationale
  4. 04
    Wireframes (low-fi)
    Sketch the main screens with layout and basic structure.
    low-fi prototyping
    what

    AI sketches low-fidelity wireframes for 5-10 main screens — gray boxes, no color.

    • 1
      Mobile-first: sketch mobile viewport first, desktop after
    • 2
      Gray rectangles + placeholder text — focus on layout + hierarchy, not aesthetics
    • 3
      Per screen annotations
      • Primary action — most important thing user does here
      • Navigation — entry/exit points
      • Content blocks — type, priority, size estimate
    why
    • ·
      Lo-fi because high-fi too early = team debates colors instead of debating flow
    • ·
      Lo-fi cheap to discard — wrong → fix in 30 minutes, not 3 days
    • ·
      Mobile-first reveals priorities: limited space forces what is actually important to surface
    how
    • ·
      Crazy 8s technique: sketch 8 variations in 8 minutes before committing
    • ·
      Gray only — no brand colors, no real images, no real text
    • ·
      Annotation mandatory — no annotation = design intent cannot be shared
  5. 05
    Interaction patterns
    Define patterns (form, search, filter, modal...) and micro-interactions.
    interaction designmicro-interactions
    what

    AI defines interaction patterns — standardize behavior across product.

    • 1
      Core patterns to define
      • Forms — validation timing (on-blur vs on-submit), error display, inline vs summary
      • Search — autocomplete, filters, results pagination
      • Modals — when to use, dismissible via ESC/backdrop, focus trap
      • Notifications — toast / inline / system / email trigger conditions
      • Loading states — skeleton / spinner / progress bar — when each
    • 2
      Each pattern: trigger condition + behavior + edge cases
    why
    • ·
      Pattern reuse = consistency — predictability builds trust
    • ·
      Define once → all forms behave same way; user learns once, applies everywhere
    • ·
      Inconsistent patterns = cognitive load on every screen the user must re-learn
    how
    • ·
      Reference Material Design / Apple HIG / Carbon for industry-standard patterns
    • ·
      Custom patterns only when the product has a strong reason — document rationale
    • ·
      Pattern library output feeds developer component specs in step 06
  6. 06
    Component specs
    Describe component types, states, variants — prepare for dev.
    atomic design
    what

    AI lists component specs — each component has states + variants + behavior.

    • 1
      Atom-level components by Atomic Design
      • Atoms — Button, Input, Label, Icon, Badge
      • Molecules — FormField, SearchBar, Card, MenuItem
      • Organisms — Header, DataTable, Modal, Sidebar
    • 2
      Each component spec
      • States — default / hover / active / disabled / error / loading
      • Variants — primary / secondary / destructive / ghost
      • Accessibility — ARIA role, keyboard, focus visible
    why
    • ·
      Component spec is the contract between UX and Dev
    • ·
      Underspec → dev guess → inconsistency across screens
    • ·
      Atomic Design (Brad Frost): build small, compose large — reusable and maintainable
    how
    • ·
      Atomic Design methodology: atoms → molecules → organisms → templates → pages
    • ·
      Each component: visual spec + behavior spec + a11y spec + code stub
    • ·
      Export to Figma component library if the team uses Figma
  7. 07
    Responsive design
    Breakpoints, mobile/tablet/desktop behaviors, adaptive patterns.
    mobile-firstprogressive enhancement
    what

    AI defines responsive behavior: breakpoints + per-breakpoint layout decisions.

    • 1
      Standard breakpoints
      • Mobile — <640px: single column, bottom nav
      • Tablet — 640-1024px: 2 column, side nav optional
      • Desktop — >1024px: multi-column, expanded nav
    • 2
      Per-breakpoint decisions
      • Navigation pattern changes (bottom nav → side nav → top nav)
      • Content priority (mobile hides secondary actions)
      • Component reflow (card stack vs grid)
    why
    • ·
      70% web traffic mobile — mobile-first is not an afterthought
    • ·
      Mobile-first vs desktop-first determines downstream code structure
    • ·
      Explicit breakpoints prevent developers from picking random values (768? 800? 1024?)
    how
    • ·
      Progressive enhancement: start mobile, add capabilities for larger screens
    • ·
      Content priority per breakpoint: what shows, what hides, what reflows
    • ·
      Test on real devices — DevTools doesn't capture touch behavior
  8. 08
    Accessibility
    WCAG 2.1 AA: keyboard nav, screen reader, contrast, focus states.
    WCAG 2.1ARIA
    what

    AI specifies accessibility per WCAG 2.1 AA — not a nice-to-have.

    • 1
      Core WCAG 2.1 AA requirements
      • Contrast — 4.5:1 for normal text, 3:1 for large text
      • Keyboard nav — all interactive elements reachable via Tab
      • Screen reader — meaningful labels, alt text, ARIA where needed
      • Focus visible — clear focus indicator on all elements
      • Error identification — errors described in text, not only color
      • Motion — prefers-reduced-motion respected
    why
    • ·
      Legal requirement in many markets: ADA (US), EAA (EU 2025)
    • ·
      Cost of retrofit 10x cost of build-in
    • ·
      Accessibility improvements help all users — SEO, keyboard power users, mobile
    how
    • ·
      WCAG 2.1 AA checklist + ARIA patterns
    • ·
      Tools: aXe DevTools, Lighthouse a11y audit
    • ·
      Manual test: keyboard-only navigation + screen reader (NVDA/VoiceOver)
  9. 09
    Design system references
    Link to the design system / tokens, or recommend creating one.
    design tokens
    what

    AI references or proposes building a design system with semantic tokens.

    • 1
      Design tokens — W3C standard
      • Color — semantic naming: color.bg.primary not color.blue.500
      • Typography — scale: xs/sm/md/lg/xl, line-heights, weights
      • Spacing — 4px base grid, named scale: 1/2/3/4/6/8/12/16...
      • Shadow, radius, motion — consistent values
    • 2
      1 token change propagates everywhere — brand refresh = update 20 tokens, not 200 screens
    why
    • ·
      Without design system: every screen reinvents wheels
    • ·
      Token-based approach lets 1 change propagate everywhere
    • ·
      Critical for Level 3-4 projects with multiple feature teams
    how
    • ·
      W3C Design Tokens standard format
    • ·
      Reference: Tailwind, Material, Carbon, Polaris — pick closest to team's stack
    • ·
      Semantic naming: describes role, not appearance
  10. 10
    Finalize spines
    Save DESIGN.md + EXPERIENCE.md, hand off to Architecture.
    specification
    what

    AI saves the two peer spines — DESIGN.md (visual identity + tokens) and EXPERIENCE.md (IA, behavior, states, a11y, journeys; references DESIGN.md tokens) — and clearly hands off to Architecture.

    • 1
      Spine structure
      • User journeys — EXPERIENCE.md, named-protagonist
      • Information Architecture — EXPERIENCE.md, sitemap + nav
      • Interaction primitives + state patterns — EXPERIENCE.md
      • Component patterns — behavioral in EXPERIENCE.md, visual specs in DESIGN.md
      • Accessibility floor — EXPERIENCE.md (visual contrast in DESIGN.md)
      • Design tokens + brand/style — DESIGN.md
    • 2
      AI invokes bmad-help → suggest Architecture (Phase 3) next
    why
    • ·
      UX spec is an input to Architecture — component list informs frontend framework choice
    • ·
      Accessibility spec informs library choice (a11y-first vs not)
    • ·
      Interaction patterns inform state management approach
    how
    • ·
      Handoff format standard BMAD — frontmatter complete with stepsCompleted
    • ·
      Wireframes: embed as ASCII/text description if Figma is not available
    • ·
      bmad-help arg='ux-design' → recommend next: architecture
IN
PRD.md
OUT
DESIGN.mdEXPERIENCE.md
WDS · 3
UX Scenarios WDS
Turns the Trigger Map into linear "sunshine path" journeys that expose every page and decision point to design. WDS Phase 3. Reads the PRD + Trigger Map.
managed byFreya 🎨 · WDS Designer
IN
trigger-map.mdPRD.md
OUT
C-UX-Scenarios/00-ux-scenarios.md
WDS · 4 · UX
UX Design WDS
The core design phase — conceptualize page structure and content direction from scenarios. Activity [UX]. WDS Phase 4. "The spec is the new code."
managed byFreya 🎨 · WDS Designer
IN
ux-scenarios
OUT
page structure direction
WDS · 4 · SP
Write Specifications WDS
Write complete development-ready page specs — content, interaction logic, spacing, typography, state variations, object registry. Activity [SP]. WDS Phase 4.
managed byFreya 🎨 · WDS Designer
IN
page structure
OUT
page spec files
WDS · 4 · SA
Spec Audit WDS
Self-audit spec completeness and quality before handoff — catches missing content, undefined interactions, empty states, SEO gaps. Activity [SA]. Run before Mimir. WDS Phase 4.
managed byFreya 🎨 · WDS Designer
IN
page spec files
OUT
validated specs
WDS · 6
Asset Generation WDS
Where specs become pixels — wireframes, page designs, UI elements, icons, images, copy via Stitch / Nano Banana. Activity [GA]. WDS Phase 6.
managed byFreya 🎨 · WDS Designer
IN
page specs + references
OUT
visual assets
WDS · 7
Design System WDSoptional
Living component library + design tokens extracted from actual page specs — grows organically, not designed upfront. Activity [DS]. WDS Phase 7 · optional.
managed byFreya 🎨 · WDS Designer
IN
page specs + assets
OUT
D-Design-System/
WDS · 5 · proto
Create UI Prototype WDSparallel after P7
After Phase 7, Mimir validates the design in code — an interactive prototype built from the specs. Runs in parallel with Design Delivery. WDS Phase 5 [P].
managed byMimir 🔨 · WDS Builder
IN
validated specs
OUT
interactive prototype
WDS · 4 · DD
Design Delivery WDSparallel after P7
Packages the validated flows into a Work Order for development handoff — Freya's [DD]/[H] activity. Hands off to Mimir in Implementation. WDS Phase 5 (handover).
managed byFreya 🎨 · WDS Designer
IN
validated specs + prototype
OUT
Work Order → Mimir
03
Solutioning
Decide how to build · Break work into stories
Required
03 · A
Architecture
Intent-routed bmad-architecture skill (create / update / validate) — lean-spine model producing ARCHITECTURE-SPINE.md as the primary source of truth, with a breadth-coverage rubric and a Reviewer Gate. Supersedes the deprecated fixed 11-step architecture flow.
managed byArchitect
  1. 01
    Detect intent & activation mode
    Create / Update / Validate — resolved from the conversation and input, not quizzed. Headless runs follow references/headless.md end to end; forwarded calls from the deprecated bmad-create-architecture shim honor pre-resolved fields verbatim.
    intent routingheadless
    what

    Resolves customization, loads bmm/config.yaml, then detects one of three intents: Create (default), Update an existing spine, or Validate one without changing it.

    why

    A single skill replacing three fixed flows must route correctly on the first turn — misrouting into Create when the user meant Validate wastes the whole session.

    how

    If a run folder already exists under {workflow.spine_output_path}, offer to resume from its memlog instead of restarting. If the real ask is requirements/UX/a capability contract/epic breakdown, redirect to bmad-prd, bmad-ux, bmad-spec, or bmad-create-epics-and-stories instead.

  2. 02
    Choose Coaching vs Fast path
    Coaching (default) — open-ended elicitation, load-bearing calls shown not silently made. Fast — draft the whole spine with [ASSUMPTION] tags for review.
    coaching pathfast path
    what

    Offered as an activation step, in the user's language, before any drafting. Coaching path pulls decisions out of the user with open-ended questions; Fast path infers and tags.

    why

    Elicitation is the value the skill is coaching toward — silently drafting the whole spine defeats the purpose unless the user explicitly wants speed.

    how

    Also asks, mandatory on both paths: is the spine the only deliverable, or does the user need a purpose-scoped human-facing artifact (team walkthrough, board vision doc) later at Finalize?

  3. 03
    Read the input to know the job
    Spec package, raw idea, sprawling doc to distill, existing codebase, one feature's slice, or an existing spine to extend/pressure-test — the input's shape decides the job.
    brownfield investigationspec-first
    what

    A SPEC.md + memlog is the richest, preferred start. Brownfield work investigates real code and project-context.md to ratify existing conventions rather than invent new ones.

    why

    The spine's altitude must mirror what it augments (initiative→features, feature→epics, epic→stories) and stay coherent with whatever level sits below it.

    how

    Inheriting a parent spine (e.g. one epic of an existing feature spine): its ADs and paradigm load as binding, read-only Inherited Invariants — only the parent's Deferred items are this run's job.

  4. 04
    Establish paradigm & seed
    Lead with a named design paradigm; recommend a verified current starter for greenfield; keep seed (stack, tree, data shape) minimal — only invariants are fixed.
    paradigmstarter recommendation
    what

    One test decides what belongs in the spine: could two units built independently choose incompatibly, and is the call non-obvious and a real trade-off? If not, it's Deferred.

    why

    Everything structural (the seed) is true at cold-start and owned by the code once it exists — over-fixing seed content invites drift between spine and reality.

    how

    Verify any named technology's current version and fit on the web before binding it into the spine.

  5. 05
    Log decisions to the memlog
    Every decision, constraint, version, assumption, and open question lands as one append-only line in .memlog.md — the run's working memory, not the rendered spine.
    memlogAD-n
    what

    Each surviving decision becomes an AD-n (stable ID, Binds / Prevents / Rule); a decision that lives only in a diagram is still logged.

    why

    Distilling the spine from a living log — instead of hand-editing the artifact — lets Update runs amend a Rule in place and add the next AD-n without ever renumbering or reusing a retired ID.

    how

    Writes go through the shared memlog.py init / append --type <decision|constraint|version|assumption|question|direction|event> script — never a hand-patch to the spine file.

  6. 06
    Distill the spine
    Write ARCHITECTURE-SPINE.md from the memlog — invariants first, seed minimal, every AD carrying Binds/Prevents/Rule, Deferred naming what it won't decide.
    breadth-coverage rubricARCHITECTURE-SPINE.md
    what

    Sweeps the breadth the altitude owns — every structural dimension (including the operational/environmental envelope: deployment, infra, operations) is decided, deferred, or an open question. No placeholders; nothing invented to fill a gap.

    why

    A whole dimension left silent is the failure mode this rubric exists to catch — not a clean spine.

    how

    A subagent per load-bearing input reconciles the draft against its source and flags anything the AD structure quietly dropped, before the Reviewer Gate.

  7. 07
    Reviewer Gate
    Deterministic lint_spine.py pass + a good-spine rubric walker + every finalize_reviewers lens dispatched as parallel subagents against ARCHITECTURE-SPINE.md.
    reviewer gatelint_spine.py
    what

    At Finalize/Update, clear fixes from the gate are applied directly. Under the standalone Validate intent, the same gate instead produces a bespoke HTML report and hands findings back to the user.

    why

    Scaled to stakes — a small feature slice gets a light pass; a platform-wide spine gets the full lens set.

    how

    Open questions and [ASSUMPTION] tags triage into blockers (resolved one at a time before handoff) and deferred items (logged with a revisit condition).

  8. 08
    Renderings, close & handoff
    Optional human-facing artifact scoped to the up-front purpose; set status: final; recommend bmad-spec, then bmad-create-epics-and-stories or bmad-create-story.
    handoffbmad-spec
    what

    The spine is the build deliverable. Optional extras — an interactive HTML+SVG walkthrough deck, a fuller solution design doc, a C4 set, a team/epic split view — are built only if the up-front purpose called for one.

    why

    AD IDs stay stable across handoff so downstream skills (epics, stories, spec) can cite them without drift.

    how

    Sets the spine's own frontmatter status: final, logs a memlog.py append --type event --text "spine finalized", then leads with bmad-spec (adopt/refresh the spine as a spec companion) before bmad-create-epics-and-stories or bmad-create-story at epic altitude.

IN
PRD.mdDESIGN.mdEXPERIENCE.md
OUT
ARCHITECTURE-SPINE.mdADRs/*.md
03 · D
Readiness Check
Gate check (v6.9) — 6-step sequential. Step 1 interactive only; Steps 2–6 auto. Validates FR coverage, UX alignment, epic quality. Verdict: READY / NEEDS WORK / NOT READY. Severity: Critical / Major / Minor.
managed byPM
  1. 01
    Load planning docs
    Read PRD, architecture, epics, UX spines (DESIGN.md + EXPERIENCE.md) — gather all planning artifacts.
    document review
    what

    PO agent loads all planning artifacts and builds inventory.

    • 1
      Load documents
      • prd.md + addendum.md + decision-log.md
      • ARCHITECTURE-SPINE.md + ADRs/*.md
      • epic-{N}.md all epics
      • DESIGN.md + EXPERIENCE.md if available
    • 2
      PO builds artifact inventory: file list + version + last-updated
    • 3
      Flag missing artifacts before starting review
    why
    • ·
      Readiness check is the gatekeeper before Phase 4
    • ·
      PO ≠ PM — independent reviewer reduces self-confirmation bias
    • ·
      BMAD tip: use different LLM for validation — same LLM finds its own writing consistent
    how
    • ·
      PO profile: skeptical, thorough, user-advocate perspective
    • ·
      Artifact inventory: date comparison (architecture newer than PRD = check for drift)
    • ·
      Missing artifacts = automatic Concerns flag
    • ·
      Version pin: record the commit/hash of each doc to ensure the review uses the correct version
  2. 02
    PRD quality review
    Check completeness, clarity, FRs/NFRs complete and non-contradictory.
    completeness check
    what

    AI reviews PRD: completeness + consistency + FR quality.

    • 1
      Completeness check
      • Every required section present and non-empty?
      • Every FR has ID + title + description + priority + user story?
      • Every NFR has a metric + threshold + measurement method?
      • Exec summary present and covering all major points?
    • 2
      Consistency check
      • Exec summary match body detail?
      • FR-IDs referenced consistently (FR-03 in one place, FR-3 elsewhere = error)?
      • NFRs do not contradict FRs?
    why
    • ·
      PRD bug most expensive when fixed early — propagates downstream to architecture and stories
    • ·
      1 missing FR = 1 missing feature = customer complaint
    • ·
      1 vague NFR = battle between dev and stakeholder post-ship
    how
    • ·
      Template-driven check: section X exists? Content X has substance?
    • ·
      ID consistency regex: normalize all references before check
    • ·
      Output: checklist with PASS/FAIL/CONCERN per item
  3. 03
    Architecture review
    ADRs have clear rationale, NFRs are addressed, risks have mitigations.
    arch review board
    what

    AI reviews Architecture: ADR quality + NFR coverage + risk register.

    • 1
      ADR quality
      • Every ADR has complete Context/Decision/Consequences/Alternatives?
      • Status field set (Proposed/Accepted)?
      • Do all tech stack decisions have ADRs?
    • 2
      NFR coverage
      • Each NFR from PRD has a component/pattern addressing it in the architecture?
      • Does the 99.9% uptime NFR have an HA topology?
      • Does the security NFR have a threat model?
    • 3
      Risk register present? Do the top 5 risks have mitigations?
    why
    • ·
      Architecture review is often skipped — 'trust the architect'
    • ·
      Independent review catches: missing ADR (decision not documented), under-addressed NFR
    • ·
      NFR-architecture gap = failure surprise in production
    how
    • ·
      ATAM-light: each NFR → find addressing component → rate adequacy
    • ·
      Tech stack alignment against team skills check
    • ·
      Vendor lock-in risk graded
  4. 04
    Cohesion check
    PRD ↔ Architecture: every FR has a solution, every decision is grounded in PRD.
    traceability matrix
    what

    AI run cohesion check: PRD ↔ Architecture traceability.

    • 1
      Build traceability matrix
      • Rows = FRs from PRD
      • Cols = architecture components
      • X = component addresses FR
    • 2
      2 types of issues
      • Empty row — FR has no component → architecture incomplete
      • Empty column — component not justified by any FR → over-engineering
    • 3
      Flag all gaps as Blocker if FR = Must priority
    why
    • ·
      Cohesion gap is a silent killer — PRD says 'users upload files', architecture has no FileService
    • ·
      Discovered after shipping = emergency fix, downtime risk
    • ·
      Over-engineering = wasted sprint capacity + maintenance burden
    how
    • ·
      Same traceability matrix technique as 03·B step 01
    • ·
      PO perspective: 'which feature will users miss?' → check corresponding FR row
    • ·
      ADR cross-check: every architecture decision traces back to a PRD requirement?
  5. 05
    Epic quality review
    Validate against best practices: user value, independence, dependencies.
    INVEST validation
    what

    AI reviews epics: INVEST compliance + AC quality + dependency graph.

    • 1
      Per story INVEST check
      • Independent? (no hard block on adjacent story?)
      • Estimable? (dev can estimate without major uncertainty?)
      • Small? (≤3 days dev work?)
      • Testable? (≥2 AC defined?)
    • 2
      AC quality check
      • Stories with 0-1 AC → insufficient spec flag
      • Stories with >7 AC → too large, split flag
      • Gherkin format: Given/When/Then present?
    • 3
      Dependency graph cycle check: A→B→A = circular dependency = blocker
    why
    • ·
      Bad stories at this gate = bad sprint — sprint failure more expensive
    • ·
      Story without clear AC → code no one can verify
    • ·
      Circular dependency → sprint never completes
    how
    • ·
      Automated INVEST scoring: count AC, check size estimate present
    • ·
      Cycle detection: graph traversal (DFS cycle check)
    • ·
      Output: per-story scorecard with specific issues
  6. 06
    Final assessment
    Conclusion: PASS, CONCERNS (minor fixes needed), or FAIL (must go back).
    go/no-gogate review
    what

    AI concludes with a binary verdict and saves the readiness report.

    • 1
      Verdict criteria
      • PASS — 0 Blockers + ≤3 Concerns → ready for Phase 4
      • NEEDS REVIEW — 0 Blockers + >3 Concerns → proceed with documented caveats
      • FAIL — ≥1 Blocker → must rework before Phase 4
    • 2
      Report structure
      • Scorecard — per-section PASS/CONCERNS/FAIL
      • Blockers — must-fix list with owner
      • Concerns — should-fix list with rationale
      • Verdict — go/no-go with reasoning
    why
    • ·
      Binary gate (go/no-go) is the critical output
    • ·
      Without explicit gate, project drifts into Phase 4 with planning gaps → sprint chaos
    • ·
      CONCERNS state can still proceed, but with documented risks — informed decision
    how
    • ·
      Blocker definition: issue that WILL cause Phase 4 failure if not fixed
    • ·
      Concern definition: issue that MIGHT cause problems but manageable
    • ·
      Save: readiness-report.md with timestamp and verdict
IN
PRD.mdARCHITECTURE-SPINE.mdepic-{N}.mdDESIGN.mdEXPERIENCE.md
OUT
readiness-report.md
03 · B
Test Design (System Level)
System-level test strategy from architecture, ADRs, and PRD — testability review (controllability, observability, reliability), risk register with P×I scoring across 6 categories, NFR thresholds, and coverage plan. Runs after architecture is complete, before the implementation-readiness gate.
operated byTest Architect (TEA)
  1. 01
    Detect mode & prerequisites
    Determine System-Level vs Epic-Level mode; confirm PRD, ADR, and architecture doc are available; halt with a clear message if required inputs are missing.
    mode detectionprerequisites
  2. 02
    Load context & knowledge base
    Extract tech stack, integration points, and NFRs from inputs; load tiered knowledge fragments just-in-time (core always, extended on-demand) for 40–50% context saving vs full load.
    tiered loadingknowledge base
  3. 03
    Testability review
    Evaluate architecture for controllability (can we set preconditions?), observability (can we verify outcomes?), and reliability (are test results deterministic?); classify ASRs as Actionable or FYI.
    controllabilityobservabilityreliability
  4. 04
    Risk assessment & prioritization
    Score 6 risk categories (TECH/SEC/PERF/DATA/BUS/OPS) with P×I matrix (1–9 scale); assign P0–P3 priority to each test scenario; high-risk items (score ≥ 6) require mandatory mitigation with owner and timeline.
    P×I scoringrisk registerP0–P3
  5. 05
    Generate output documents
    Write test-design-architecture.md (testability concerns + risk mitigations for Dev/Arch teams), test-design-qa.md (coverage plan + execution recipe for QA team), and BMAD handoff document bridging to epic decomposition.
    test-design-architecture.mdBMAD handoff
IN
ARCHITECTURE-SPINE.mdADRs/*.mdPRD.md
OUT
test-design-architecture.mdtest-design-qa.mdtest-design/{project}-handoff.md
03 · E
Framework Setup
Initialize a production-ready test framework — directory structure, config, fixtures (mergeTests pattern), Faker data factories, sample tests, and README. Fully autonomous: auto-detects stack (frontend/backend/fullstack) and selects the appropriate framework with no manual scaffolding. Resumable if interrupted.
operated byTest Architect (TEA)
  1. 01
    Preflight checks
    Read config, auto-detect stack from project manifests (package.json, pyproject.toml, pom.xml, go.mod); verify no conflicting framework exists; confirm write permissions; save detected context to progress file.
    stack detectionprerequisites
  2. 02
    Framework selection
    Frontend: default to Playwright, switch to Cypress only if criteria met. Backend: Python→pytest, Java/Kotlin→JUnit 5, Go→go test, C#→xUnit, Ruby→RSpec, Rust→cargo test. Announce selection with rationale.
    PlaywrightCypresspytest
  3. 03
    Scaffold framework
    Generate complete directory tree, framework config, .env.example, fixtures with mergeTests pattern (auto-cleanup hooks built in), Faker-based data factories, sample tests, and helpers. Adaptive execution: parallel agent-team or sequential based on runtime capability.
    mergeTests patternfixturesFaker factories
  4. 04
    Documentation & scripts
    Create tests/README.md covering setup, running, debugging, architecture, best practices, and CI notes; add idiomatic test commands to package.json, Makefile, or pyproject.toml (minimum: "test:e2e" for frontend).
    tests/README.mdnpm scripts
  5. 05
    Validate & summarize
    Check all checklist items (directory structure, config correctness, fixtures/factories, docs/scripts); fix any gaps before marking complete; report framework selected, all artifacts created, and recommended next steps.
    validation checklistcompletion report
IN
package.jsonpyproject.toml / pom.xml / go.mod
OUT
framework config filetests/ directory treetests/README.md
03 · C
Epics & Stories
Created AFTER architecture (v6.9) — stories reflect technical decisions correctly. 4-step micro-file: validate prerequisites → design epics → generate stories → final validation. UX-DRs are first-class inputs. JIT entity creation.
managed byPM + Architect
  1. 01
    Load PRD + architecture
    Read both to understand requirements and existing technical decisions.
    traceability
    what

    AI loads PRD + ARCHITECTURE-SPINE.md and builds FR ↔ component traceability map.

    • 1
      AI loads both documents in parallel
      • prd.md — FRs, NFRs, personas, success metrics
      • ARCHITECTURE-SPINE.md — components, data model, API design, patterns
    • 2
      Build traceability matrix
      • Rows = FRs from PRD
      • Cols = architecture components
      • X = 'this component addresses this FR'
    • 3
      Flag gaps: FR has no component (architecture incomplete) + component has no FR (over-engineering)
    why
    • ·
      v6 improvement: stories are created AFTER architecture — stories now reference specific patterns
    • ·
      Stories created before architecture miss tech context → rework mid-sprint
    • ·
      Traceability surfaces gaps BEFORE sprint capacity is committed
    how
    • ·
      Traceability matrix: rows = FRs, cols = components, X = mapping
    • ·
      Empty row = FR uncovered → must address in architecture
    • ·
      Empty col = component not justified → over-engineering, remove or justify
  2. 02
    Group FRs into epics
    Group related FRs into epics — each epic carries independent value.
    capability mapping
    what

    AI groups FRs into epics — each epic is a user-deliverable capability.

    • 1
      AI clusters FRs by user capability, not by technical layer
    • 2
      Validate per epic
      • User-deliverable: can it be demonstrated standalone?
      • Independent value: do users benefit if only this epic ships?
      • Reasonable scope: 2-6 weeks development?
    • 3
      Bad vs good epic naming
      • Bad: 'Backend', 'Frontend' — technical layers are not capabilities
      • Good: 'User onboarding', 'Payment processing', 'Admin dashboard'
    why
    • ·
      Tech-layer epics do not deliver value individually
    • ·
      User-centric epics deliver value per epic ship — stakeholders see progress
    • ·
      Validate 'demo standalone' as the acid test for epic independence
    how
    • ·
      Capability mapping: each FR maps to exactly 1 epic
    • ·
      Epic size check: one epic too large (>6 weeks) → split; too small (<1 week) → merge
    • ·
      Validate against the product vision: each epic must tell a capability story
  3. 03
    Epic ordering
    Determine epic order by dependencies, risk, and value.
    WSJFvalue stream
    what

    AI orders epics by WSJF (Weighted Shortest Job First).

    • 1
      WSJF scoring per epic
      • User-Business Value — direct revenue/UX impact (1-10)
      • Time Criticality — cost of delaying (1-10)
      • Risk Reduction — reduces unknowns if done early (1-10)
      • Job Size — estimated weeks (1-10, higher = larger)
      • WSJF = (Value + Criticality + Risk) / Size
    • 2
      High-risk, high-value epic → ship early (fail fast while resources remain)
    • 3
      First epic must be independent (no dependencies) — unblock parallel work
    why
    • ·
      Shipping in the wrong order = waste — first 4 sprints does not deliver demoable value = stakeholder frustration
    • ·
      Risky epics early: if they fail → pivot while still have runway
    • ·
      Value-first: demonstrate value early → stakeholder confidence + feedback
    how
    • ·
      SAFe WSJF framework
    • ·
      Score subjectively if exact data is unavailable — relative ranking is enough
    • ·
      Visualize: table with scores + bar chart so stakeholders can see the rationale
  4. 04
    Story breakdown
    Split epics into stories using INVEST principles.
    INVESTSPIDR
    what

    AI splits each epic into INVEST stories.

    • 1
      INVEST checklist per story
      • Independent — do not block each other when possible
      • Negotiable — scope can be adjusted
      • Valuable — clear user/business benefit
      • Estimable — effort can be estimated
      • Small — 1-3 days of dev work is ideal
      • Testable — AC can be verified
    • 2
      SPIDR splitting patterns when the story is too large
      • Spike — separate exploration/research story
      • Path — split by user journey paths
      • Interface — split by UI vs API vs integration
      • Data — split by data variations
      • Rules — split by business rules
    why
    • ·
      Story too large → uncertainty → wrong estimates → sprint slips
    • ·
      Story too small → overhead (PR review, deploy) > value
    • ·
      INVEST sweet spot: 1-3 days reduces risk and enables fast feedback
    how
    • ·
      INVEST checklist per story — fail any = must split/reshape
    • ·
      SPIDR patterns: most used = Interface split (API vs UI separate stories)
    • ·
      Story size: count by AC — story with >7 AC is usually too large, split
  5. 05
    Acceptance criteria
    Write AC in Given/When/Then format for each story, with enough detail to test.
    BDDGherkin
    what

    AI writes Acceptance Criteria in BDD/Gherkin format — executable specification.

    • 1
      Gherkin format
      • Given — precondition (state before the action)
      • When — action user takes
      • Then — expected outcome (verifiable)
    • 2
      Coverage per story
      • 1+ AC for happy path
      • 1-2 edge cases (boundary values, empty states)
      • 1 error path (what happens when it fails)
    • 3
      3-7 AC per story — under 3 = insufficient spec; over 7 = story too large
    why
    • ·
      Vague AC ('works correctly') → endless debate about done
    • ·
      BDD format forces precondition + trigger + outcome — do not miss edge cases
    • ·
      Executable AC: if using Cucumber/Playwright, tests write themselves from AC
    how
    • ·
      Each AC independently testable — do not chain multiple outcomes into one AC
    • ·
      Edge cases from INVEST size estimate: 'what are the boundaries of this story?'
    • ·
      Error paths mandatory: 'What does the user see when it fails?'
  6. 06
    Story dependencies
    Identify dependencies between stories and mark blockers.
    dependency graph
    what

    AI builds a dependency graph between stories.

    • 1
      3 dependency types
      • Hard blocker — story B cannot start until A is done
      • Soft dependency — B is easier if A is done, but B can still start
      • Independent — no dependency
    • 2
      AI identifies the critical path — longest blocker chain
    • 3
      Flag stories with >3 dependencies — signal must split or re-sequence
    why
    • ·
      Dependency graph drives sprint planning and parallel work
    • ·
      Critical path = minimum sprint duration — visualize it for the SM
    • ·
      Unresolved dependencies = sprint blocker surprise
    how
    • ·
      Build graph: nodes = stories, directed edges = blocker
    • ·
      Critical path analysis: longest path = minimum sprint duration
    • ·
      Parallel opportunities: stories with no outgoing dependencies = safe to parallelize
    • ·
      Cross-epic dependency = red flag: story depending on another epic breaks sprint boundaries
  7. 07
    Finalize epic files
    Save epic-{N}.md for each epic with all stories inside.
    backlog
    what

    AI save epic-{N}.md per epic — stories nested inside.

    • 1
      File structure
      • 1 file per epic: epic-01-user-onboarding.md
      • Frontmatter: epic ID, status, owner, WSJF score
      • Stories inside: S1, S2, S3... with full AC and technical context
    • 2
      Sharding for large epics
      • Option: stories/epic-01/story-01.md separated
      • Use when epic has >10 stories — avoid one huge file
    • 3
      AI invokes bmad-help → suggest Readiness Check (03·D) next
    why
    • ·
      1 file 1 epic = just-in-time loading — Phase 4 loads only 1 epic at a time
    • ·
      Full story content in the epic file = dev agent reads only the epic, no PRD/architecture required
    • ·
      Separation ensures minimal context window per sprint
    how
    • ·
      Frontmatter: epicId, status, owner, sequence, wsjfScore
    • ·
      Stories numbered S1, S2 within epic — globally unique via epic-01-S1
    • ·
      bmad-help recommend: Readiness Check (03·D) before Phase 4
IN
PRD.mdARCHITECTURE-SPINE.md
OUT
epic-{N}.mdstories/*.md
03 · F
CI/CD Pipeline
Scaffold a production-ready CI quality pipeline — parallel sharding (4 runners, fail-fast: false), burn-in flaky detection (10× iterations), dependency caching, artifact collection on failure, and quality gate enforcement (P0=100%, P1≥95%). Supports GitHub Actions, GitLab CI, Jenkins, Azure DevOps, Harness CI, and Circle CI. Prerequisite: test framework must be set up first and local tests must pass.
operated byTest Architect (TEA)
  1. 01
    Preflight checks
    Verify Git repo exists, detect stack, confirm test framework config is present, run local tests (halt with "Fix failing tests before CI setup" if they fail); detect CI platform from existing config files or git remote URL; default to GitHub Actions.
    platform detectionprerequisites
  2. 02
    Generate CI pipeline
    Select platform-specific template; configure all stages (lint → test → contract-test → burn-in → report); set up parallel sharding (4 shards via matrix strategy); route all user-controlled contexts through env: intermediaries to prevent script injection.
    parallel shardingpipeline stagesinjection prevention
  3. 03
    Configure quality gates & notifications
    Enable burn-in loop (10× iterations, frontend/fullstack by default; backend skipped as deterministic); set quality thresholds (P0=100%, P1≥95%); configure can-i-deploy contract gate if Pact is enabled; wire Slack/email failure notifications with artifact links.
    burn-in loopquality gatesP0/P1 thresholds
  4. 04
    Validate & summarize
    Validate generated config against the full checklist (config path, stages, sharding, burn-in, artifacts, secrets); fix any gaps before marking complete; write completion report with config path, enabled stages, required secrets, and next steps (set secrets, push, run pipeline).
    secrets checklistvalidation
IN
framework config filespassing local tests
OUT
CI config filescripts/test-changed.shscripts/ci-local.shdocs/ci.mddocs/ci-secrets-checklist.md
04
Implementation
BMAD governs story/status and drives Dev Story execution · TEA owns pre-code acceptance scaffolds and post-code release evidence · WDS/Mimir runs Create UI Prototype in parallel as a design-validation step
Required
BMAD / BMM
Sprint planning, Create Story, sprint status, course correction, retrospectives, and investigations remain the governance layer. The story file is canonical; prompt-only sync/status steps are not cards.
WDS / Mimir
Mimir's only footprint in this phase is the parallel Create UI Prototype workflow — an interactive prototype built in code from Freya's validated specs, run alongside Design Delivery to validate the design before the production build. BMAD Dev Story remains the implementation driver for every story, WDS-enabled or not.
TEA
TEA runs Test Design and ATDD before Dev Story execution, then Automation, Test Review, NFR Assessment, and Traceability after Dev Story verification.
Integration rule
BMAD Dev Story implements every story — WDS-spec or not; there is no separate Mimir implementation track. Mimir's Create UI Prototype runs in parallel as a design-validation step, not a competing execution track. Both remain gated by TEA and governed by BMAD.
upstream
BMAD Context
PRD · Arch · Epics
once per sprint
Sprint Plan
BMAD
per story
Create Story
BMAD
pre-code strategy
Test Design
TEA
pre-code gate
ATDD
TEA
per story
Dev Story
BMAD
after verify
Automate + Review
TEA · QA
release evidence
NFR + Trace
TEA · QA
closure
Gate / Status / Retro
Prompt · BMAD
ATDD rule: ATDD is a sequential pre-code gate after Create Story and before BMAD Dev Story implementation starts. Dev Story consumes the ATDD scaffold/checklist directly; the developer activates each relevant skipped test, observes red, implements the behavior, and turns it green. The remaining TEA evidence work resumes only after Amelia's code review and Dev Story completion produce fresh implementation evidence.
04 · A
Sprint Planning BMAD
Initialize and refresh sprint-status.yaml from epics and story files before any implementation cycle starts.
run byAmelia / Dev
  1. 01
    Parse epic files
    Scan epic files and extract epic numbers plus story IDs/titles from headers in the body.
    epic discovery
    what

    Scan epic files and extract epic numbers plus story IDs/titles from headers in the body.

    why

    · Sprint planning needs ALL epics and stories to build complete status tracking

    how

    · Whole doc takes priority over sharded docs when both exist

  2. 02
    Build status structure
    Create a flat development_status map with default statuses for epics, stories, and retrospectives.
    flat map default status
    what

    Create a flat development_status map with default statuses for epics, stories, and retrospectives.

    why

    Create a flat development_status map with default statuses for epics, stories, and retrospectives.

    how

    · Order: epic → that epic's stories → retrospective → next epic

  3. 03
    Apply status detection
    Check real story files, upgrade status when files exist, and never downgrade.
    file detection never downgrade
    what

    Check real story files, upgrade status when files exist, and never downgrade.

    why

    Check real story files, upgrade status when files exist, and never downgrade.

    how

    · Check path: {story_location}/{story-key}.md

  4. 04
    Generate sprint-status.yaml
    Write the flat development_status map with metadata header to the file.
    artifact
    what

    Write the flat development_status map with metadata header to the file.

    why

    Write the flat development_status map with metadata header to the file.

    how

    · Output path: {implementation_artifacts}/sprint-status.yaml

  5. 05
    Validate + Report
    Check completeness and report totals before finishing.
    validation
    what

    Check completeness and report totals before finishing.

    why

    · Artifacts drive state — YAML persists across sessions, every agent reads the same source

    how

    Validate completeness and report totals. Validate that every epic/story/retro has an entry, there are no extra entries, status values are valid, and YAML syntax is correct

IN
epic filesstory filesimplementation_artifacts
OUT
sprint-status.yaml
04 · C
Create Story BMAD
Create the ready-for-dev BMAD story just in time from sprint status, epic context, architecture, and prior learnings.
run byAmelia / Dev
  1. 01
    Determine target story
    Use a provided story path or auto-discover the first backlog story from sprint-status.yaml.
    story selection
    what

    Read the full sprint-status.yaml, select the first story entry with status backlog, or parse the user-provided story id/path.

    why

    Create Story is just-in-time per story. Preparing stories in bulk makes later stories stale after earlier implementation learnings.

    how

    If this is the first story in an epic, move the epic from backlog to in-progress. Halt if no eligible backlog story exists.

  2. 02
    Load and analyze core artifacts
    Load epics, PRD, architecture, UX, previous story learnings, and recent git patterns selectively.
    artifact synthesisprevious story intelligence
    what

    Extract the target story's user statement, ACs, dependencies, constraints, and any context from prior story files or recent commits.

    why

    The story file must contain everything downstream agents need without rereading full planning docs.

    how

    Load epics in full; load PRD, architecture, and UX sections selectively for the chosen story.

  3. 03
    Extract architecture guardrails
    Capture tech stack, file locations, naming conventions, API contracts, data schemas, security, and regression constraints.
    architecture guardrails
    what

    Read every UPDATE file the story is expected to touch and document current behavior, intended change, and what must be preserved.

    why

    Wrong file locations, stale APIs, and missed existing behavior are the common failure modes Create Story prevents.

    how

    Make the guardrails explicit in Dev Notes so BMAD Dev Story inherits them directly during implementation.

  4. 04
    Research latest technical specifics
    Check current library versions, changelogs, security notes, deprecations, and best practices relevant to the story.
    web researchfreshness
    what

    Research only the APIs, frameworks, and external services the selected story will interact with.

    why

    LLM memory may be stale; story creation is the last cheap point to catch breaking changes or vulnerable patterns.

    how

    Embed specific versions, API notes, migration cautions, and security patches in the story file.

  5. 05
    Create comprehensive story file
    Write the complete story file from template and set Status = ready-for-dev.
    story fileready-for-dev
    what

    Synthesize requirements, ACs, dev notes, architecture compliance, test expectations, and current-tech notes into {story_key}.md.

    why

    The output becomes the canonical story input for Dev Story and every downstream implementation gate.

    how

    Initialize from template, fill every required section, and save clarification questions for the end instead of interrupting the workflow.

  6. 06
    Update sprint status and finalize
    Validate the story file, update sprint-status.yaml to ready-for-dev, preserve comments, and report the handoff.
    status updatehandoff
    what

    Run checklist validation and update only the matching story status plus last_updated in sprint-status.yaml.

    why

    ready-for-dev means the story is prepared for Dev Story but implementation has not started.

    how

    Report story id, path, status, and next step: Dev Story implementation.

IN
sprint-status.yamlepic contextarchitectureprevious story notes
OUT
{story_key}.mdsprint-status.yaml
04 · B
Epic Level Test Design TEA
Create or refresh the epic-level risk and coverage architecture used by the story and ATDD before implementation starts.
run byMurat / TEA
  1. 01
    Load story and epic context
    Read the selected story, parent epic, architecture, project context, and approved spec.
    contextFR inventory
    what

    Read the selected story, parent epic, architecture, project context, and approved spec.

    why

    Coverage must be anchored to actual requirements and architecture constraints.

    how

    Inventory functional requirements, scenarios, NFRs, and integration boundaries.

  2. 02
    Assess risk
    Score probability and impact for each scenario or requirement group.
    risk
    what

    Score probability and impact for each scenario or requirement group.

    why

    The highest-risk behavior deserves the strongest evidence.

    how

    Map risks into P0-P3 priorities and record the rationale.

  3. 03
    Choose test levels
    Assign unit, integration, contract, E2E, and NFR checks to the right risks.
    test pyramid
    what

    Assign unit, integration, contract, E2E, and NFR checks to the right risks.

    why

    A balanced pyramid gives confidence without creating a brittle suite.

    how

    Use unit tests for logic, integration for boundaries, E2E for critical journeys, and NFR checks for quality attributes.

  4. 04
    Generate coverage plan
    Write the per-story coverage strategy and evidence expectations.
    coverage plan
    what

    Write the per-story coverage strategy and evidence expectations.

    why

    The plan becomes the test architecture contract for ATDD, automation, and release traceability.

    how

    Document test levels, fixtures, data, priority, and owner.

  5. 05
    Validate and hand off
    Check for uncovered or untestable requirements and hand gaps to ATDD, Dev, or PM.
    handoff
    what

    Check for uncovered or untestable requirements and hand gaps to ATDD, Dev, or PM.

    why

    Gaps found here are cheaper than gaps found at release gate.

    how

    Flag missing ACs, ambiguous requirements, and coverage decisions needing human approval.

IN
epic-{N}.mdprd.mdARCHITECTURE-SPINE.mdproject-context.md
OUT
test-design-epic-{N}.md
04 · D
ATDD TEA
Generate skipped red-phase acceptance scaffolds before production code. Dev activates each relevant test, observes red, then turns it green through TDD.
run byMurat / TEA
  1. 01
    Read acceptance criteria
    Extract behaviors, examples, error paths, and observable outcomes.
    AC
    what

    Extract behaviors, examples, error paths, and observable outcomes.

    why

    ATDD must test user-visible behavior, not implementation guesses.

    how

    Map each AC to one or more executable examples.

  2. 02
    Write behavior scenarios
    Translate acceptance criteria into Given/When/Then or equivalent spec cases.
    scenario
    what

    Translate acceptance criteria into Given/When/Then or equivalent spec cases.

    why

    Scenario language exposes missing preconditions and expected outcomes.

    how

    Name scenarios after behavior, not functions.

  3. 03
    Generate skipped scaffold
    Create skipped/disabled acceptance tests, fixtures, and helper stubs without implementing production behavior.
    red scaffold
    what

    Create skipped/disabled acceptance tests, fixtures, and helper stubs without implementing production behavior.

    why

    The scaffold defines the expected red signal without letting implementation start before ATDD handoff is complete.

    how

    Keep assertions meaningful, mark each scaffold as intentionally skipped/disabled, and record the command that should turn red after activation.

  4. 04
    Prepare red checklist
    Record which scaffolds Dev must activate, the expected red failure, and the command to run.
    handoff checklist
    what

    Record which scaffolds Dev must activate, the expected red failure, and the command to run.

    why

    ATDD owns acceptance intent and scaffold quality; Dev owns the red observation after activation inside BMAD Dev Story's TDD loop.

    how

    Do not require Murat to turn tests red here. The checklist tells Dev Story exactly what to activate and what red signal to expect.

  5. 05
    Hand off to Dev Story
    Attach scaffold paths, activation checklist, test commands, and expected red failures to the story file for Dev Story.
    handoff
    what

    Attach scaffold paths, activation checklist, test commands, and expected red failures to the story file for Dev Story.

    why

    Dev Story must consume ATDD output before implementation starts so coding starts from an acceptance contract.

    how

    Link tests to tasks and acceptance criteria, then pass the checklist directly into the story file Dev Story reads.

IN
synced storyapproved spec.mdacceptance criteriatest-design-epic-{N}.md
OUT
tests/**/*.spec.tsatdd-checklist-{story_key}.md
04 · DEV
Dev Story BMAD
Implement the story via TDD — auto-discover the ready-for-dev story, RED-GREEN-REFACTOR each task without stopping, validate the Definition of Done, hand off for Code Review.
run byAmelia / Dev
  1. 01
    Find & load story
    Auto-discover the story with status=ready-for-dev in sprint-status.yaml, or use the story path the user provides; read the FULL story file.
    story discoverysprint-status
    what

    Determine which story to implement and read its entire content.

    • ·
      User provides a story path → use it directly
    • ·
      No path given → read the FULL sprint-status.yaml → find the first N-N-name entry with status ready-for-dev
    • ·
      Read the COMPLETE story file: AC, Tasks/Subtasks, Dev Notes, Dev Agent Record, File List, Change Log, Status
    • ·
      Load project-context.md if it exists
    • ·
      HALT if no story is found or the story file is not accessible
    why
    • ·
      The story file is the single source of truth — every implementation decision comes from it
    • ·
      Dev Notes carry architecture guardrails and learnings from Create Story — skipping them loses critical context
    how
    • ·
      Read the FULL sprint-status.yaml (every line) → filter N-N-name entries → take the first with status=ready-for-dev
    • ·
      If none is ready-for-dev → ask the user to run Create Story, validate it, or provide a story path
  2. 02
    Detect review continuation
    Check whether the story already went through Code Review — if so, extract pending follow-up items to process first.
    review continuation
    what
    • ·
      Scan the story file: does it have a "Senior Developer Review (AI)" section?
    • ·
      If yes → review_continuation = true: extract the review outcome, count unchecked [AI-Review] items, save the list of pending items
    • ·
      If no → review_continuation = false: fresh start
    why
    • ·
      Review follow-ups must be processed before regular tasks — the reviewer already identified specific issues
    • ·
      Skipping detection risks silently skipping high-severity review findings
    how
    • ·
      Scan the "Review Follow-ups (AI)" subsection under Tasks/Subtasks
    • ·
      Count High / Med / Low severity items still unchecked
  3. 03
    Mark in-progress
    Capture baseline_commit into the story's YAML frontmatter; update sprint-status.yaml: ready-for-dev → in-progress.
    baseline commitstate tracking
    what
    • ·
      Only run if the current status = ready-for-dev and the story YAML has no baseline_commit yet
    • ·
      Run git rev-parse HEAD → save into the story's YAML frontmatter as baseline_commit: <sha>
    • ·
      Update sprint-status.yaml: story_key → in-progress, update last_updated
    why
    • ·
      baseline_commit tells Code Review which commit the diff should start from
    • ·
      Status in-progress signals to the team that the story is actively being implemented
    how
    • ·
      If git is not available → baseline_commit: NO_VCS
    • ·
      If the YAML frontmatter already has baseline_commit → do not overwrite it
    • ·
      Update sprint-status.yaml: rewrite the FULL file, preserving ALL comments and structure
  4. 04
    Implement task — RED·GREEN·REFACTOR loop
    Repeat per task: write failing tests (RED) → implement minimal code (GREEN) → refactor → validate → mark [x] → next task, without stopping midway.
    TDDred-green-refactorzero-interrupt
    what

    Loop through each task in the order listed in the story's Tasks/Subtasks:

    • ·
      RED: write failing tests → confirm tests fail before writing code
    • ·
      GREEN: implement the minimal code needed to pass — nothing outside task scope
    • ·
      REFACTOR: improve code structure; tests stay green after refactor
    • ·
      Validate: run the full test suite, check the AC tied to the task
    • ·
      Mark the task [x] only once all its tests actually exist and pass 100%
    • ·
      If review_continuation = true: process the [AI-Review] items before regular tasks
    why
    • ·
      TDD forces the behavior to be articulated before an implementation is chosen — avoids over-engineering
    • ·
      NEVER stop for "milestones," "significant progress," or "session boundaries" — only HALT on a specific blocking event
    • ·
      NEVER mark [x] unless tests actually pass — no lying or cheating about test status
    how
    • ·
      HALT if: a dependency outside the story spec is needed, 3 consecutive implementation failures occur, or required config is missing
    • ·
      Update the File List after each task: ALL new/modified/deleted files (relative paths)
    • ·
      Record completion notes in the Dev Agent Record after each task completes
  5. 05
    Story completion & DoD validation
    Verify all tasks are [x], run full regression, validate the Definition of Done, update status: in-progress → review.
    definition of doneregression
    what

    Gate before the story enters the review queue:

    • ·
      Re-scan the entire story document: verify ALL tasks and subtasks are already [x]
    • ·
      Run the full regression suite — don't skip it
    • ·
      DoD checklist: all AC satisfied, unit/integration/e2e tests added, linting passes, File List complete, Dev Agent Record has notes, Change Log has a summary
    • ·
      Update story Status → review
    • ·
      Update sprint-status.yaml: story_key → review, preserving ALL comments
    why
    • ·
      DoD validation before signaling review avoids wasting the reviewer's time on incomplete stories
    • ·
      Status review (not done) — the story waits for Code Review approval before it's done
    how
    • ·
      HALT if: any task is not yet [x], regression fails, File List is missing, or DoD fails
    • ·
      YAML update: change only status and last_updated — don't rewrite the whole file
  6. 06
    Communication & handoff
    Report completion tailored to user_skill_level; suggest running Code Review with a different LLM.
    handoff
    what
    • ·
      Summary: story ID, story key, title, key changes, tests added, files modified
    • ·
      Provide the story file path and current status (review)
    • ·
      Ask whether the user needs anything explained (per user_skill_level)
    • ·
      Suggest next steps: run Code Review, verify ACs, check sprint status
    why
    • ·
      The user needs to know the story is done and may have questions about implementation decisions
    • ·
      Code Review should use a different LLM to avoid blind spots from the model that already implemented the story
    how
    • ·
      Tailor the explanation to user_skill_level — junior needs more detail, senior needs trade-offs
    • ·
      Tip: "For best results, run Code Review using a different LLM than the one that implemented this story"
IN
sprint-status.yaml{story_key}.mdproject-context.md
OUT
source code + testssprint-status.yaml
04 · REV
Code Review BMAD
Gather context via a 5-tier cascade → spawn 3 parallel adversarial reviewers → triage findings → log into the story file, resolve, and update story status.
run byAmelia / Dev (collab TEA)
  1. 01
    Gather context
    Find the review target via a 5-tier cascade; construct the diff; determine review_mode: full (spec/story file present) or no-spec.
    5-tier cascadereview mode
    what

    Find the review target via a 5-tier cascade (stop at the first tier that matches):

    • ·
      Tier 1: explicit arg — PR, commit SHA, branch, spec file, diff-mode keyword (staged/uncommitted/branch diff/commit range)
    • ·
      Tier 2: recent conversation context — scan the most recent messages
    • ·
      Tier 3: sprint-status.yaml → find a story with status=review
    • ·
      Tier 4: current git state — if branch ≠ main, confirm with the user
    • ·
      Tier 5: ask the user directly
    • ·
      Set review_mode: full (spec/story file present) or no-spec
    why
    • ·
      The cascade avoids asking the user unnecessarily — whatever is found automatically is used automatically
    • ·
      review_mode=full → Acceptance Auditor runs; no-spec → Acceptance Auditor is skipped
    how
    • ·
      Construct the diff: staged / uncommitted / branch diff / commit range / provided diff — verify it's non-empty before continuing
    • ·
      If the diff is > ~3000 lines → warn the user, offer chunking
    • ·
      HALT + present a summary (diff stats, review_mode, loaded docs) for the user to confirm before reviewing
  2. 02
    Parallel review — 3 subagents
    Spawn 3 adversarial reviewer subagents in parallel, each receiving different context.
    parallel reviewmulti-perspective
    what

    Dev spawns 3 parallel adversarial reviewer subagents, each receiving separate context:

    • ·
      Blind Hunter: receives the diff ONLY — no spec, no project context. Finds logic errors, security issues, runtime exceptions. Doesn't know the intended behavior → catches unintended ones
    • ·
      Edge Case Hunter: receives the diff + project read access. Traces all branching paths. Finds boundary errors, race conditions, null/empty/overflow cases
    • ·
      Acceptance Auditor: receives the diff + spec file + context docs. Only runs when review_mode=full. Checks whether each AC is implemented, AC drift, missing behaviors
    why
    • ·
      The 3 perspectives are orthogonal — Blind Hunter: logic bugs; Edge Case: path gaps; Auditor: AC drift — no overlap, no misses
    • ·
      A single reviewer misses at least one dimension — multi-agent review gives better coverage
    how
    • ·
      All 3 run simultaneously (parallel), not sequentially
    • ·
      If subagents aren't available → generate prompt files, HALT; the user runs each file in a separate session (ideally a different LLM) and pastes findings back
    • ·
      If a layer fails or returns empty → log it into failed_layers, continue with findings from the remaining layers
  3. 03
    Triage findings
    Normalize → deduplicate → classify into 4 action buckets: decision_needed, patch, defer, dismiss.
    normalizededuplicateclassify
    what

    Automatically process findings from the 3 layers:

    • ·
      Normalize: convert the 3 different output formats (Blind Hunter: markdown; Edge Case: JSON; Auditor: markdown) into a unified format: id, source, title, detail, location
    • ·
      Deduplicate: an issue flagged by 2-3 reviewers is merged into one, with source set to e.g. "blind+edge"
    • ·
      Classify each finding into exactly one bucket: decision_needed (ambiguous, needs a human), patch (unambiguous fix), defer (pre-existing, not caused by this change), dismiss (noise/false positive)
    • ·
      Drop all dismiss findings — log the count in the report
    why
    • ·
      Without dedup, the human must sift through 3x the noise
    • ·
      4 action buckets are clearer than severity labels — each bucket has a specific action
    • ·
      Zero findings after triage → "Clean review" — without forcing a minimum number of findings
    how
    • ·
      If review_mode=no-spec: any decision_needed finding is reclassified → patch (if the fix is clear) or defer
    • ·
      If failed_layers is non-empty: warn the user before announcing the result
  4. 04
    Present, resolve & update status
    Record findings into the story file; HALT for the user to resolve decision_needed items and patches; update story status and sprint-status.yaml.
    story file updatestatus update
    what
    • ·
      Record findings into the story file: append a "Review Findings" subsection under Tasks/Subtasks in the format - [ ] [Review][Decision/Patch/Defer] title
    • ·
      Record defer items into deferred-work.md
    • ·
      HALT: resolve decision_needed findings — the user decides (patch, defer, or dismiss)
    • ·
      HALT: process patch findings — the user chooses: apply all / leave as action items / walk through each
    • ·
      Update story status: all resolved → done; action items remain → in-progress
    • ·
      Update sprint-status.yaml: story_key → new status, preserving ALL comments
    why
    • ·
      Writing into the story file means findings persist in the source of truth — nothing is lost when the session ends
    • ·
      decision_needed findings must have human input — AI cannot decide ambiguous choices on its own
    • ·
      Status done is set only once all issues are resolved — the story is never left in limbo
    how
    • ·
      If no-spec: no story file exists → present findings in chat only, don't persist them
    • ·
      If zero findings after triage → "Clean review" → skip straight to status update
    • ·
      Suggest next: run Dev Story (next story) or re-run Code Review after fixes
IN
git diff (required)story-[slug].md (full mode)sprint-status.yaml (required)
OUT
story-[slug].md (findings appended)deferred-work.mdsprint-status.yaml
04 · E
Automate TEA
After implementation exists, expand automation beyond ATDD happy paths into edge cases, negative paths, regressions, and critical journeys.
run byMurat / TEA
  1. 01
    Inspect existing coverage
    Read current tests, coverage reports, and changed files.
    coverage
    what

    Read current tests, coverage reports, and changed files.

    why

    Automation should fill real gaps, not duplicate existing evidence.

    how

    Map gaps to risk priorities from Test Design.

  2. 02
    Select automation targets
    Choose edge cases, regressions, integration points, and critical user journeys.
    targeting
    what

    Choose edge cases, regressions, integration points, and critical user journeys.

    why

    Not every path deserves E2E coverage.

    how

    Match test level to risk and maintenance cost.

  3. 03
    Generate tests
    Add automated tests with realistic fixtures and assertions.
    generate
    what

    Add automated tests with realistic fixtures and assertions.

    why

    Useful automation checks behavior and fails clearly.

    how

    Avoid brittle selectors and over-mocked integration behavior.

  4. 04
    Run and stabilize
    Run the suite repeatedly enough to catch flakiness.
    stabilize
    what

    Run the suite repeatedly enough to catch flakiness.

    why

    Flaky tests erode trust in CI.

    how

    Fix timing, data, and isolation issues before handoff.

  5. 05
    Update CI notes
    Document commands, test grouping, and CI configuration changes.
    CI
    what

    Document commands, test grouping, and CI configuration changes.

    why

    Automation only helps if the team can run it consistently.

    how

    Add config changes or explicit follow-up items.

IN
implemented codeATDD outputscoverage reports
OUT
expanded test suiteCI notes
04 · F
QA · Test Review TEA QA
Audit test quality after code and automation exist, scoring determinism, isolation, performance, and maintainability.
run byQA / TEA
  1. 01
    Load Context & Knowledge Base
    Determine review scope, detect stack, and load tiered knowledge base.
    scope
    what

    Determine review scope (single / directory / suite), detect stack, load knowledge fragments, gather optional context (story file, test-design doc, framework config).

    why

    Quality evaluation rules are stack-conditional — frontend, backend, and fullstack have different patterns and anti-patterns.

    how

    Scan playwright.config.*, pom.xml, go.mod etc. Core knowledge always loaded; Playwright Utils and Pact.js Utils loaded on demand based on stack and config flags.

  2. 02
    Discover & Parse Tests
    Collect test files matching scope, parse metadata, and optionally capture browser evidence.
    discover
    what

    Collect test files, parse metadata per file (framework, describe/test counts, fixtures, factories, network interception, waits, control flow), optionally capture browser trace and screenshots.

    why

    Structural analysis catches anti-patterns that static reading misses — flaky waits, shared mutable state, conditional logic inside tests.

    how

    Halt if no tests found. Use Playwright CLI trace analysis for failing assertion root-cause when browser automation is enabled.

  3. 03
    Orchestrate Quality Evaluation (4 Workers)
    Run 4 parallel quality workers — Determinism, Isolation, Performance, Maintainability — then aggregate results.
    evaluate
    what

    Dispatch all 4 quality workers in parallel; each writes a JSON result. Step 3F aggregates: weighted overall score, top-10 prioritised recommendations.

    why

    Each dimension requires different expertise — parallel workers are 60–70% faster than sequential and prevent one dimension from biasing another.

    how

    Resolve mode: agent-team → subagent → sequential (capability probe). Wait for all 4 outputs before aggregating.

  4. 04
    Generate Report & Validate
    Produce test-review.md with overall score, dimension breakdown, critical findings, and recommended next workflow.
    report
    what

    Produce test-review.md: overall score + grade, dimension breakdown, critical findings with fixes, warnings, context references, and recommended next workflow.

    why

    A scored, graded report gives the team a concrete quality baseline and clear action items — not just a list of problems.

    how

    Validate against checklist.md before finalising. Note: test-review scores quality, not coverage — direct coverage gaps to bmad-testarch-trace.

IN
test filesframework configstory/test-design context
OUT
test-review.mdquality score
04 · G
NFR Assessment TEA QA
Assess security, performance, reliability, and scalability evidence before Traceability computes the release gate state.
run byQA / TEA
  1. 01
    Load Context & Knowledge Base
    Check prerequisites, load knowledge base, and gather NFR thresholds from tech-spec, PRD, or story.
    context
    what

    Confirm implementation is accessible and evidence sources are available. Load knowledge base and context artifacts — NFR thresholds from tech-spec.md → PRD.md → story/test-design (priority order).

    why

    NFR evaluation without thresholds is guesswork — HALT if prerequisites are missing rather than proceeding with incomplete data.

    how

    Summarise which sources were found and what evidence is available before moving to Step 2.

  2. 02
    Define NFR Categories & Thresholds
    Select the 8 standard ADR categories, extract thresholds, and build the NFR matrix.
    thresholds
    what

    8 standard categories: Testability · Data Strategy · Scalability · Disaster Recovery · Security · Observability · Quality of Service · Deployability. Extract thresholds per category; mark UNKNOWN if not found.

    why

    Any UNKNOWN threshold means that category reports as CONCERNS — missing thresholds are never silently skipped.

    how

    Never guess a threshold. Pull from sources in priority order: tech-spec → PRD → story/test-design.

  3. 03
    Gather Evidence
    Collect measurable data per category — performance metrics, security reports, error history, DR drill records.
    evidence
    what

    Scan for: P95 response times, throughput, OWASP ZAP / Snyk reports, error rate history, burn-in results, DR drill records. Optionally capture browser network data via Playwright CLI.

    why

    Missing evidence is never "no issue" — any category without concrete evidence is automatically marked CONCERNS.

    how

    Store browser capture outputs in {test_artifacts}/nfr/. Close session after capture.

  4. 04
    Orchestrate NFR Evaluation (4 Workers)
    Run 4 parallel domain workers — Security, Performance, Reliability, Scalability — then aggregate with worst-case rule.
    evaluate
    what

    Dispatch Workers A–D in parallel: Security (04a), Performance (04b), Reliability (04c), Scalability (04d). Each outputs risk level (NONE / LOW / MEDIUM / HIGH) and PASS / CONCERN / FAIL per sub-category.

    why

    Parallel workers are independent — one domain's risk does not bias another's evaluation.

    how

    Sub-step 4E aggregates all 4 outputs. Worst-case rule: overall risk = highest risk among 4 workers. Detect cross-domain compound risks (e.g. Performance + Scalability).

  5. 05
    Generate Report & Validate
    Produce nfr-assessment.md with per-category results, overall risk level, remediation actions, and compliance matrix.
    report
    what

    Write nfr-assessment.md: per-category results with evidence summary, overall risk level, remediation actions (URGENT / HIGH / MEDIUM / LOW), compliance matrix (SOC2 / GDPR / HIPAA / PCI-DSS), CI gate YAML snippet.

    why

    A risk-rated report with remediation urgency lets the team prioritise fixes before the release gate — not after.

    how

    Validate against checklist.md. Recommend next workflow: run bmad-testarch-trace for coverage gate, or proceed to release if NFR is PASS.

IN
implemented codeNFR thresholdsmetrics/scans/logs
OUT
nfr-assessment.md
04 · H
Traceability TEA
Map requirements to evidence and produce the final coverage gate after Automation, Test Review, and NFR Assessment are available.
run byMurat / TEA
  1. 01
    Inventory requirements
    List functional requirements, acceptance criteria, and NFRs.
    requirements
    what

    List functional requirements, acceptance criteria, and NFRs.

    why

    Traceability starts with a complete requirement set.

    how

    Mark ambiguous or changed requirements explicitly.

  2. 02
    Map tests to requirements
    Connect each requirement to one or more test cases and result evidence.
    mapping
    what

    Connect each requirement to one or more test cases and result evidence.

    why

    Coverage must be visible at requirement level.

    how

    Include test file, case name, level, and latest result.

  3. 03
    Identify gaps
    Flag uncovered, weakly covered, flaky, or waived requirements.
    gaps
    what

    Flag uncovered, weakly covered, flaky, or waived requirements.

    why

    Gaps are release risks and need owner decisions.

    how

    Classify by severity and recommended action.

  4. 04
    Produce gate verdict
    Summarize PASS, CONCERNS, FAIL, or WAIVED.
    gate
    what

    Summarize PASS, CONCERNS, FAIL, or WAIVED.

    why

    The team needs a clear release posture.

    how

    Require rationale and owner for every concern or waiver.

  5. 05
    Hand off actions
    Send required fixes to automation, review, Correct Course, or PM.
    handoff
    what

    Send required fixes to automation, review, Correct Course, or PM.

    why

    Traceability should trigger action, not just reporting.

    how

    Link actions to requirement IDs and test evidence.

IN
story/ACstest resultsnfr-assessment.mdtest-review.md
OUT
traceability matrixPASS/CONCERNS/FAIL/WAIVED gate
04 · I
Correct Course BMAD
Off-ramp when the gate, implementation, or stakeholder evidence reveals material scope or architecture change.
run byPM / Dev
  1. 01
    Initialize
    Confirm change trigger; ask user description; verify access to PRD + Epics (required); choose mode: Incremental / Batch.
    change triggermode selection
    what

    Confirm change trigger; ask user description; verify access to PRD + Epics (required); choose mode: Incremental / Batch.

    why

    · Trigger unclear → analysis misdirected; HALT early saves wasted effort

    how

    · Ask numbered questions; if trigger still unclear after answer response → HALT and require specific details

  2. 02
    Analyze change
    Follow checklist.md — 6 sections: trigger/context, epic impact, artifact conflicts, path forward evaluation, proposal components, final review. Record [x] / [N/A] / [!] each item.
    checklistimpact analysis
    what

    Follow checklist.md — 6 sections: trigger/context, epic impact, artifact conflicts, path forward evaluation, proposal components, final review. Record [x] / [N/A] / [!] each item.

    why

    · Checklist systematic = does not miss impacts; ad-hoc analysis often misses downstream effects

    how

    · Section by section, present progress after each major section

  3. 03
    Draft proposals
    Create explicit edit proposals per artifact — old → new format + rationale. Incremental: present each one; Batch: collect all, present together.
    edit proposalsold → new
    what

    Create explicit edit proposals per artifact — old → new format + rationale. Incremental: present each one; Batch: collect all, present together.

    why

    · Old → new format = unambiguous, not misunderstanding about exactly what change

    how

    · Story format: Story: [ID] Title / Section: ... / OLD: ... / NEW: ... / Rationale: ...

  4. 04
    Generate proposal
    Compile Sprint Change Proposal document: 5 sections → save to sprint-change-proposal-{date}.md . Present to user, ask Continue [c] / Edit [e].
    Sprint Change Proposaldocument
    what

    Compile Sprint Change Proposal document: 5 sections → save to sprint-change-proposal-{date}.md . Present to user, ask Continue [c] / Edit [e].

    why

    · Structured document = stakeholder has specific review, approve, route that does not need re-explain context

    how

    · Save to {planning_artifacts}/sprint-change-proposal-{date}.md

  5. 05
    Finalize & route
    Get explicit user approval (yes/no/revise) → classify scope (Minor/Moderate/Major) → route: Minor→Dev / Moderate→PO+Dev / Major→PM+Architect. Update sprint-status.yaml.
    approvalhandoffrouting
    what

    Get explicit user approval (yes/no/revise) → classify scope (Minor/Moderate/Major) → route: Minor→Dev / Moderate→PO+Dev / Major→PM+Architect. Update sprint-status.yaml.

    why

    · Explicit approval = not ambiguity about "did we decide this?" — prevents team from acting on unapproved changes

    how

    · Minor → Dev agent: deliverables = finalized edit proposals + implementation tasks

  6. 06
    Complete
    Summary: issue addressed, scope classification, artifacts modified, routed to. Confirm deliverables and next steps.
    completionsummary
    what

    Summary: issue addressed, scope classification, artifacts modified, routed to. Confirm deliverables and next steps.

    why

    · Explicit close = clear transition point; team knows workflow ended and who does what next

    how

    · Personalized completion message to user; list every artifact already modified or created

IN
gate concernschange triggerevidence
OUT
change proposalhandoff decision
04 · SS
Sprint Status BMAD
Inspect and update sprint-status.yaml after prompt-only status transitions such as review to done.
run byAmelia / Dev
  1. 01
    Locate file
    Find sprint-status.yaml; if it does not exist → exit with guidance to run sprint-planning.
    file lookup exit if missing
    what

    Find sprint-status.yaml; if it does not exist → exit with guidance to run sprint-planning.

    why

    Find sprint-status.yaml; if it does not exist → exit with guidance to run sprint-planning.

    how

    Find sprint-status.yaml; if it does not exist → exit with guidance to run sprint-planning.

  2. 02
    Parse & analyze
    Read full sprint-status.yaml, classify keys (epics: epic-* / retros: *-retrospective / stories: remaining), count statuses (backlog/ready-for-dev/in-progress/review/done), detect risks (stale >7d, orphan story, in-progress epic no stories, story in review).
    classify count statuses detect risks
    what

    Read full sprint-status.yaml, classify keys (epics: epic-* / retros: *-retrospective / stories: remaining), count statuses (backlog/ready-for-dev/in-progress/review/done), detect risks (stale >7d, orphan story, in-progress epic no stories, story in review).

    why

    Read full sprint-status.yaml, classify keys (epics: epic-* / retros: *-retrospective / stories: remaining), count statuses (backlog/ready-for-dev/in-progress/review/done), detect risks (stale >7d, orphan story, in-progress epic no stories, story in review).

    how

    Read full sprint-status.yaml, classify keys (epics: epic-* / retros: *-retrospective / stories: remaining), count statuses (backlog/ready-for-dev/in-progress/review/done), detect risks (stale >7d, orphan story, in-progress epic no stories, story in review).

  3. 03
    Select recommendation
    Priority: in-progress story → review → ready-for-dev → backlog → retrospective → all done.
    priority order next action
    what

    Priority: in-progress story → review → ready-for-dev → backlog → retrospective → all done.

    why

    Priority: in-progress story → review → ready-for-dev → backlog → retrospective → all done.

    how

    Priority: in-progress story → review → ready-for-dev → backlog → retrospective → all done.

  4. 04
    Display summary
    Project info + story/epic counts + next recommendation + risk list.
    summary risks
    what

    Project info + story/epic counts + next recommendation + risk list.

    why

    Project info + story/epic counts + next recommendation + risk list.

    how

    Project info + story/epic counts + next recommendation + risk list.

  5. 05
    Offer actions
    [1] run recommended workflow / [2] show stories by status / [3] show raw yaml / [4] exit.
    interactive user choice
    what

    [1] run recommended workflow / [2] show stories by status / [3] show raw yaml / [4] exit.

    why

    [1] run recommended workflow / [2] show stories by status / [3] show raw yaml / [4] exit.

    how

    [1] run recommended workflow / [2] show stories by status / [3] show raw yaml / [4] exit.

IN
sprint-status.yamlstory status
OUT
status summarynext actionopen action_items
04 · J
Retrospective BMAD
At end of epic, capture learning from Dev Story implementation and TEA traceability evidence before the next sprint planning loop.
run byAmelia / Dev
  1. 01
    Epic discovery
    Find the recently completed epic from sprint-status; confirm with the user; verify all stories are done.
    sprint-statusepic verification
    what

    Find the recently completed epic from sprint-status; confirm with the user; verify all stories are done.

    why

    · Retro on epic not yet done = missing data → insights insufficient

    how

    · If partial: the user chooses whether to continue or HALT to finish stories first

  2. 02
    Deep story analysis
    Read EACH story file in the epic — extract dev notes, review feedback, lessons, tech debt, testing insights → synthesize cross-story patterns.
    story analysispattern synthesis
    what

    Read EACH story file in the epic — extract dev notes, review feedback, lessons, tech debt, testing insights → synthesize cross-story patterns.

    why

    · Story files contain real data — rich material for discussion, not memory-based assumptions

    how

    · Read each file: {implementation_artifacts}/{N}-{M}-*.md

  3. 03
    Previous retro integration
    Load retro epic before — check action item follow-through, lessons were applied or skipped.
    continuityaccountability
    what

    Load retro epic before — check action item follow-through, lessons were applied or skipped.

    why

    · Not checking the previous retro = repeating the same mistakes and losing accountability

    how

    · If this is epic 1 (no previous retro): set first_retrospective = true and skip this step

  4. 04
    Next epic preview
    Load epic {N+1} — identify dependencies, gaps, preparation needs before start.
    dependency mappingpreparation
    what

    Load epic {N+1} — identify dependencies, gaps, preparation needs before start.

    why

    · Prep work identified during the retro → start epic {N+1} with solid foundation instead of hit blockers mid-epic

    how

    · If epic {N+1} is not yet defined: skip, but still proceed with the retro normally

  5. 05
    Epic review discussion
    Party mode: Amelia facilitates what went well / what didn't / patterns — INTERACTIVE with user, blameless, system-focused.
    party modepsychological safety
    what

    Party mode: Amelia facilitates what went well / what didn't / patterns — INTERACTIVE with user, blameless, system-focused.

    why

    · Party mode = psychological safety — everyone speaks honestly, not blame, not defensive

    how

    · Ground rules: no blame, focus on systems, specific examples preferred, every voice counts

  6. 06
    Next epic prep + action items
    Synthesize action items; plan next epic preparation tasks; significant change detection; critical readiness check.
    action itemssignificant change detection
    what

    Synthesize action items; plan next epic preparation tasks; significant change detection; critical readiness check.

    why

    · Significant change detection prevents starting epic {N+1} on wrong assumptions → mid-epic failure

    how

    · Detect significant change: architectural assumptions wrong, scope change, tech approach must change, newly discovered dependencies, security/compliance issues

  7. 07
    Save & update status
    Save epic-{N}-retro-{date}.md ; update sprint-status: epic-{N}-retrospective → done .
    institutional memorysprint-status
    what

    Save epic-{N}-retro-{date}.md ; update sprint-status: epic-{N}-retrospective → done .

    why

    · Date-stamped file = append-only institutional memory — next retro load file this to check follow-through

    how

    · If the epic-{N}-retrospective key is not found in sprint-status → warn the user; manual update is needed

IN
completed epic storiesDev Story completion notesTEA traceability matrix
OUT
epic-{N}-retro-{date}.mdsprint-status.yaml (action_items)
04 · —
Dev Auto BMAD
One iteration of an unattended development loop — clarify & route, plan, implement, review. Invoked by name; not yet registered in the menu-chained skill order.
run byAmelia (Dev)
  1. 01
    Clarify and route
    Clarify the requested change and route to the appropriate execution path.
    clarifyroute
  2. 02
    Plan
    Produce an implementation plan for the clarified change.
    plan
  3. 03
    Implement
    Implement the planned change.
    implement
  4. 04
    Review
    Review the implementation and finalize the iteration.
    review
IN
spec file
OUT
implementation_artifacts (spec file)bmad-dev-auto-result-*.md
WDS · 8
Product Evolution WDS
The full WDS pipeline compressed into a kaizen loop for living products — Analyze → Scope → Design → Implement → Test → Deploy, every change on its own branch. Activities A/S/D/I/T/P. Can feed new user types or driving forces back into the Trigger Map. WDS Phase 8.
managed byFreya 🎨 (+ Mimir 🔨 for implementation)
IN
existing product
OUT
scoped change on feature branch
Changelog