- Finds relevant files using glob patterns
- Finds relevant code using grep searches
- Reads and analyses file contents
- Maps dependencies between files
- Identifies existing coding patterns
- Flags risks and concerns
- Flags open questions for the human approval gate
- Produces structured exploration report
- Adjusts thoroughness by complexity level
- Uses only read-only tools (read, glob, grep — no write, edit, bash)
- Synthesises exploration report into solution design
- Considers alternative approaches
- Breaks work into phases for moderate/complex tasks
- Produces atomic task specifications with all 7 fields:
- Scope and non-goals
- Acceptance criteria (testable conditions)
- Definition of done (completion checklist)
- Automated tests
- Manual test steps
- Rollback note
- Risk level
- Produces task briefs for the Coder (scope, constraints, files, assumptions, patterns)
- Defines do-not-touch list and dependency guardrails
- Orders tasks by dependency within each phase
- Scales planning depth by complexity (simple: concise, moderate: full, complex: comprehensive)
- Each task is independently verifiable and produces a committable change
- Uses only read-only tools (read, glob, grep — no write, edit, bash)
- Reviews task brief before writing code
- Reviews task specification (acceptance criteria, definition of done)
- Reads every file before modifying it
- Verifies files are NOT on the do-not-touch list before editing
- Follows implementation plan step order
- Uses Edit tool for modifications (not Write for existing files)
- Matches existing code patterns
- Checks guardrails (no unrelated changes, no new deps without justification)
- Documents any deviations from task spec
- Reports acceptance criteria status in implementation report
- Detects bug-fixing loop (same file edited 3+ times for same issue)
- Leaves no incomplete implementations
- Reviews every changed file against the task specification
- Applies Layer 1 — Automated checks (tests, lint, type check, build)
- Applies Layer 2 — Behavioural checks (manual test steps, edge cases, failure paths)
- Applies Layer 3 — Operational checks (error handling, logging, config, migrations, rollback)
- Applies Layer 4 — Security checks (input validation, encoding, auth, secrets, hygiene, deps)
- Scales verification depth by risk level:
- Low risk: Layer 1 full, Layers 2-4 basic
- Medium risk: all layers standard
- High risk: all layers thorough with evidence
- Checks each acceptance criterion (pass/fail with evidence)
- Checks each definition of done item (complete/incomplete)
- Produces clear PASS / FAIL / PASS_WITH_WARNINGS verdict
- Provides specific, actionable fix instructions on FAIL
- Uses only bash + read tools (no write/edit)
- Refuses to start on a dirty git tree (clean-tree precondition)
- Happy path: build → verify PASS → commits with
task-{n} {description}message - Fix-round path: FAIL → fix instructions passed verbatim → re-verify (max 3 fix rounds)
- Exhaustion path: fourth FAIL →
git stash→ skip → next task starts on a clean tree - PASS_WITH_WARNINGS commits (FAIL never does)
- Never pushes; delivery ends at local commits
- Final batch report includes verdicts, attempts, commit hashes, stash refs, and failure details
- Workers return reports in the
implementing-tasks/verifying-changesformats
- Human can classify request complexity (simple / moderate / complex)
- Human can invoke
/exploreand receive a complete exploration report - Human can review exploration findings and decide on solution direction
- Human can invoke
/planand receive atomic task specifications - Human can review the plan and approve before coding
- Human can invoke
/codeper task and receive implementation reports - Human can invoke
/verifyper task and receive verification reports - Human can invoke
/commit-taskto commit verified changes - Human can manage the task loop (advance to next task)
- Human can manage the phase loop (return to
/planfor next phase) - Human can handle FAIL verdicts (re-run
/codewith fix instructions, max 2 retries) - Human can detect bug-fixing loops and re-run
/explorefor new evidence - Human can escalate when retries are exhausted
- Simple task completes all stages successfully
- Moderate task completes all stages with multiple tasks
- Complex task completes all stages with multiple phases
- Context flows correctly between commands (via conversation)
- Human Gate #1 pauses workflow for exploration review
- Human Gate #2 pauses workflow for plan review
- Each task produces a self-contained commit
- Task loop correctly advances through tasks in a phase
- Phase loop correctly returns to
/planfor next phase
- Verification FAIL triggers human to re-run
/codewith fix instructions - Coder receives specific fix instructions from Verifier
- Re-verification runs after fix
- Second retry works if first fix is insufficient
- Bug-fixing loop escape triggers after 2 retries exhausted
- Bug-fixing loop escape triggers if same file patched 3+ times
- Escape returns to
/explorefor new evidence - Escalation occurs if escape also fails
- User approves at Gate #1 → proceeds to
/plan - User rejects at Gate #1 → re-runs
/exploreor modifies direction - User approves at Gate #2 → proceeds to
/code - User rejects at Gate #2 → modifies plan or re-explores
- Simple tasks get concise approval summaries
- Complex tasks get full reports at approval gates
-
/exploreruns exploration only and delivers report -
/planproduces atomic task specs and delivers plan -
/codeimplements a single atomic task and delivers report -
/verifyruns 4-layer verification and delivers report -
/commit-taskcommits changes with task-based message -
/epcvshows workflow reference
- Task spec with all 7 fields is produced and consumed correctly
- Task brief is produced and consumed correctly
- Acceptance criteria are checked by Verifier (pass/fail per criterion)
- Definition of done is checked by Verifier (complete/incomplete per item)
- Do-not-touch list is respected by Coder
- Rollback note is verified by Verifier
- Risk level determines verification depth
- Request with no matching files in codebase
- Request that requires creating entirely new files
- Request that conflicts with existing architecture
- Request with ambiguous requirements (open questions flagged)
- Request that touches security-sensitive code (high risk verification)
- Request that requires changes to test files
- Very large codebase with many matching files
- Request that the Planner determines is infeasible
- Multi-phase project with phase loop
- Task that fails verification and triggers bug-fixing loop escape
- User rejects at both approval gates in sequence
- Start with a simple, well-defined task
- Verify each command produces expected output format
- Verify Human Gate #1 pauses and presents findings correctly
- Verify Human Gate #2 pauses and presents plan correctly
- Check that context passes correctly between commands
- Verify atomic task specs have all 7 fields
- Verify 4-layer verification report is produced
- Verify commit is created per task
- Try a moderate task and verify task loop with multiple tasks
- Try a complex task and verify phase loop
- Intentionally create a verification failure to test retry flow
- Exhaust retries to test bug-fixing loop escape
- Test each slash command independently
- Review all output formats match the report templates in the skills (
skills/implementing-tasks/SKILL.md,skills/verifying-changes/SKILL.md)