2026-02-01 16:47:50 +09:00
/**
* Prometheus Plan Template
*
* The markdown template structure for work plans generated by Prometheus.
* Includes TL;DR, context, objectives, verification strategy, TODOs, and success criteria.
*/
export const PROMETHEUS_PLAN_TEMPLATE = ` ## Plan Structure
Generate plan to: \` .sisyphus/plans/{name}.md \`
\` \` \` markdown
# {Plan Title}
## TL;DR
> **Quick Summary**: [1-2 sentences capturing the core objective and approach]
>
> **Deliverables**: [Bullet list of concrete outputs]
> - [Output 1]
> - [Output 2]
>
> **Estimated Effort**: [Quick | Short | Medium | Large | XL]
> **Parallel Execution**: [YES - N waves | NO - sequential]
> **Critical Path**: [Task X → Task Y → Task Z]
---
## Context
### Original Request
[User's initial description]
### Interview Summary
**Key Discussions**:
- [Point 1]: [User's decision/preference]
- [Point 2]: [Agreed approach]
**Research Findings**:
- [Finding 1]: [Implication]
- [Finding 2]: [Recommendation]
### Metis Review
**Identified Gaps** (addressed):
- [Gap 1]: [How resolved]
- [Gap 2]: [How resolved]
---
## Work Objectives
### Core Objective
[1-2 sentences: what we're achieving]
### Concrete Deliverables
- [Exact file/endpoint/feature]
### Definition of Done
- [ ] [Verifiable condition with command]
### Must Have
- [Non-negotiable requirement]
### Must NOT Have (Guardrails)
- [Explicit exclusion from Metis review]
- [AI slop pattern to avoid]
- [Scope boundary]
---
## Verification Strategy (MANDATORY)
2026-02-02 14:18:01 +09:00
> **UNIVERSAL RULE: ZERO HUMAN INTERVENTION**
>
> ALL tasks in this plan MUST be verifiable WITHOUT any human action.
> This is NOT conditional — it applies to EVERY task, regardless of test strategy.
>
> **FORBIDDEN** — acceptance criteria that require:
> - "User manually tests..." / "사용자가 직접 테스트..."
> - "User visually confirms..." / "사용자가 눈으로 확인..."
> - "User interacts with..." / "사용자가 직접 조작..."
> - "Ask user to verify..." / "사용자에게 확인 요청..."
> - ANY step where a human must perform an action
>
> **ALL verification is executed by the agent** using tools (Playwright, interactive_bash, curl, etc.). No exceptions.
2026-02-01 16:47:50 +09:00
### Test Decision
- **Infrastructure exists**: [YES/NO]
2026-02-02 14:18:01 +09:00
- **Automated tests**: [TDD / Tests-after / None]
2026-02-01 16:47:50 +09:00
- **Framework**: [bun test / vitest / jest / pytest / none]
### If TDD Enabled
Each TODO follows RED-GREEN-REFACTOR:
**Task Structure:**
1. **RED**: Write failing test first
- Test file: \` [path].test.ts \`
- Test command: \` bun test [file] \`
- Expected: FAIL (test exists, implementation doesn't)
2. **GREEN**: Implement minimum code to pass
- Command: \` bun test [file] \`
- Expected: PASS
3. **REFACTOR**: Clean up while keeping green
- Command: \` bun test [file] \`
- Expected: PASS (still)
**Test Setup Task (if infrastructure doesn't exist):**
- [ ] 0. Setup Test Infrastructure
- Install: \` bun add -d [test-framework] \`
- Config: Create \` [config-file] \`
- Verify: \` bun test --help \` → shows help
- Example: Create \` src/__tests__/example.test.ts \`
- Verify: \` bun test \` → 1 test passes
2026-02-02 14:18:01 +09:00
### Agent-Executed QA Scenarios (MANDATORY — ALL tasks)
2026-02-01 16:47:50 +09:00
2026-02-02 14:18:01 +09:00
> Whether TDD is enabled or not, EVERY task MUST include Agent-Executed QA Scenarios.
> - **With TDD**: QA scenarios complement unit tests at integration/E2E level
> - **Without TDD**: QA scenarios are the PRIMARY verification method
2026-02-01 16:47:50 +09:00
>
2026-02-02 14:18:01 +09:00
> These describe how the executing agent DIRECTLY verifies the deliverable
> by running it — opening browsers, executing commands, sending API requests.
> The agent performs what a human tester would do, but automated via tools.
**Verification Tool by Deliverable Type:**
2026-02-01 16:47:50 +09:00
2026-02-02 14:18:01 +09:00
| Type | Tool | How Agent Verifies |
|------|------|-------------------|
| **Frontend/UI** | Playwright (playwright skill) | Navigate, interact, assert DOM, screenshot |
| **TUI/CLI** | interactive_bash (tmux) | Run command, send keystrokes, validate output |
| **API/Backend** | Bash (curl/httpie) | Send requests, parse responses, assert fields |
| **Library/Module** | Bash (bun/node REPL) | Import, call functions, compare output |
| **Config/Infra** | Bash (shell commands) | Apply config, run state checks, validate |
2026-02-01 16:47:50 +09:00
2026-02-02 14:18:01 +09:00
**Each Scenario MUST Follow This Format:**
2026-02-01 16:47:50 +09:00
2026-02-02 14:18:01 +09:00
\` \` \`
Scenario: [Descriptive name — what user action/flow is being verified]
Tool: [Playwright / interactive_bash / Bash]
Preconditions: [What must be true before this scenario runs]
Steps:
1. [Exact action with specific selector/command/endpoint]
2. [Next action with expected intermediate state]
3. [Assertion with exact expected value]
Expected Result: [Concrete, observable outcome]
Failure Indicators: [What would indicate failure]
Evidence: [Screenshot path / output capture / response body path]
\` \` \`
2026-02-01 16:47:50 +09:00
2026-02-02 14:18:01 +09:00
**Scenario Detail Requirements:**
- **Selectors**: Specific CSS selectors ( \` .login-button \` , not "the login button")
- **Data**: Concrete test data ( \` "test@example.com" \` , not \` "[email]" \` )
- **Assertions**: Exact values ( \` text contains "Welcome back" \` , not "verify it works")
- **Timing**: Include wait conditions where relevant ( \` Wait for .dashboard (timeout: 10s) \` )
- **Negative Scenarios**: At least ONE failure/error scenario per feature
- **Evidence Paths**: Specific file paths ( \` .sisyphus/evidence/task-N-scenario-name.png \` )
**Anti-patterns (NEVER write scenarios like this):**
- ❌ "Verify the login page works correctly"
- ❌ "Check that the API returns the right data"
- ❌ "Test the form validation"
- ❌ "User opens browser and confirms..."
**Write scenarios like this instead:**
- ✅ \` Navigate to /login → Fill input[name="email"] with "test@example.com" → Fill input[name="password"] with "Pass123!" → Click button[type="submit"] → Wait for /dashboard → Assert h1 contains "Welcome" \`
- ✅ \` POST /api/users {"name":"Test","email":"new@test.com"} → Assert status 201 → Assert response.id is UUID → GET /api/users/{id} → Assert name equals "Test" \`
- ✅ \` Run ./cli --config test.yaml → Wait for "Loaded" in stdout → Send "q" → Assert exit code 0 → Assert stdout contains "Goodbye" \`
**Evidence Requirements:**
- Screenshots: \` .sisyphus/evidence/ \` for all UI verifications
- Terminal output: Captured for CLI/TUI verifications
- Response bodies: Saved for API verifications
- All evidence referenced by specific file path in acceptance criteria
2026-02-01 16:47:50 +09:00
---
## Execution Strategy
### Parallel Execution Waves
> Maximize throughput by grouping independent tasks into parallel waves.
> Each wave completes before the next begins.
2026-02-16 15:14:25 +09:00
> Target: 5-8 tasks per wave. Fewer than 3 per wave (except final) = under-splitting.
2026-02-01 16:47:50 +09:00
\` \` \`
2026-02-16 15:14:25 +09:00
Wave 1 (Start Immediately — foundation + scaffolding):
├── Task 1: Project scaffolding + config [quick]
├── Task 2: Design system tokens [quick]
├── Task 3: Type definitions [quick]
├── Task 4: Schema definitions [quick]
├── Task 5: Storage interface + in-memory impl [quick]
├── Task 6: Auth middleware [quick]
└── Task 7: Client module [quick]
Wave 2 (After Wave 1 — core modules, MAX PARALLEL):
├── Task 8: Core business logic (depends: 3, 5, 7) [deep]
├── Task 9: API endpoints (depends: 4, 5) [unspecified-high]
├── Task 10: Secondary storage impl (depends: 5) [unspecified-high]
├── Task 11: Retry/fallback logic (depends: 8) [deep]
├── Task 12: UI layout + navigation (depends: 2) [visual-engineering]
├── Task 13: API client + hooks (depends: 4) [quick]
└── Task 14: Telemetry middleware (depends: 5, 10) [unspecified-high]
Wave 3 (After Wave 2 — integration + UI):
├── Task 15: Main route combining modules (depends: 6, 11, 14) [deep]
├── Task 16: UI data visualization (depends: 12, 13) [visual-engineering]
├── Task 17: Deployment config A (depends: 15) [quick]
├── Task 18: Deployment config B (depends: 15) [quick]
├── Task 19: Deployment config C (depends: 15) [quick]
└── Task 20: UI request log + build (depends: 16) [visual-engineering]
Wave 4 (After Wave 3 — verification):
├── Task 21: Integration tests (depends: 15) [deep]
├── Task 22: UI QA - Playwright (depends: 20) [unspecified-high]
├── Task 23: E2E QA (depends: 21) [deep]
└── Task 24: Git cleanup + tagging (depends: 21) [git]
2026-02-16 15:19:31 +09:00
Wave FINAL (After ALL tasks — independent review, 4 parallel):
├── Task F1: Plan compliance audit (oracle)
├── Task F2: Code quality review (unspecified-high)
├── Task F3: Real manual QA (unspecified-high)
└── Task F4: Scope fidelity check (deep)
Critical Path: Task 1 → Task 5 → Task 8 → Task 11 → Task 15 → Task 21 → F1-F4
2026-02-16 15:14:25 +09:00
Parallel Speedup: ~70% faster than sequential
Max Concurrent: 7 (Waves 1 & 2)
2026-02-01 16:47:50 +09:00
\` \` \`
2026-02-16 15:14:25 +09:00
### Dependency Matrix (abbreviated — show ALL tasks in your generated plan)
| Task | Depends On | Blocks | Wave |
|------|------------|--------|------|
| 1-7 | — | 8-14 | 1 |
| 8 | 3, 5, 7 | 11, 15 | 2 |
| 11 | 8 | 15 | 2 |
| 14 | 5, 10 | 15 | 2 |
| 15 | 6, 11, 14 | 17-19, 21 | 3 |
| 21 | 15 | 23, 24 | 4 |
2026-02-01 16:47:50 +09:00
2026-02-16 15:14:25 +09:00
> This is abbreviated for reference. YOUR generated plan must include the FULL matrix for ALL tasks.
2026-02-01 16:47:50 +09:00
### Agent Dispatch Summary
2026-02-16 15:14:25 +09:00
| Wave | # Parallel | Tasks → Agent Category |
|------|------------|----------------------|
| 1 | **7** | T1-T4 → \` quick \` , T5 → \` quick \` , T6 → \` quick \` , T7 → \` quick \` |
| 2 | **7** | T8 → \` deep \` , T9 → \` unspecified-high \` , T10 → \` unspecified-high \` , T11 → \` deep \` , T12 → \` visual-engineering \` , T13 → \` quick \` , T14 → \` unspecified-high \` |
| 3 | **6** | T15 → \` deep \` , T16 → \` visual-engineering \` , T17-T19 → \` quick \` , T20 → \` visual-engineering \` |
| 4 | **4** | T21 → \` deep \` , T22 → \` unspecified-high \` , T23 → \` deep \` , T24 → \` git \` |
2026-02-16 15:19:31 +09:00
| FINAL | **4** | F1 → \` oracle \` , F2 → \` unspecified-high \` , F3 → \` unspecified-high \` , F4 → \` deep \` |
2026-02-01 16:47:50 +09:00
---
## TODOs
> Implementation + Test = ONE Task. Never separate.
2026-02-16 15:19:31 +09:00
> EVERY task MUST have: Recommended Agent Profile + Parallelization info + QA Scenarios.
> **A task WITHOUT QA Scenarios is INCOMPLETE. No exceptions.**
2026-02-01 16:47:50 +09:00
- [ ] 1. [Task Title]
**What to do**:
- [Clear implementation steps]
- [Test cases to cover]
**Must NOT do**:
- [Specific exclusions from guardrails]
**Recommended Agent Profile**:
> Select category + skills based on task domain. Justify each choice.
- **Category**: \` [visual-engineering | ultrabrain | artistry | quick | unspecified-low | unspecified-high | writing] \`
- Reason: [Why this category fits the task domain]
- **Skills**: [ \` skill-1 \` , \` skill-2 \` ]
- \` skill-1 \` : [Why needed - domain overlap explanation]
- \` skill-2 \` : [Why needed - domain overlap explanation]
- **Skills Evaluated but Omitted**:
- \` omitted-skill \` : [Why domain doesn't overlap]
**Parallelization**:
- **Can Run In Parallel**: YES | NO
- **Parallel Group**: Wave N (with Tasks X, Y) | Sequential
- **Blocks**: [Tasks that depend on this task completing]
- **Blocked By**: [Tasks this depends on] | None (can start immediately)
**References** (CRITICAL - Be Exhaustive):
> The executor has NO context from your interview. References are their ONLY guide.
> Each reference must answer: "What should I look at and WHY?"
**Pattern References** (existing code to follow):
- \` src/services/auth.ts:45-78 \` - Authentication flow pattern (JWT creation, refresh token handling)
**API/Type References** (contracts to implement against):
- \` src/types/user.ts:UserDTO \` - Response shape for user endpoints
**Test References** (testing patterns to follow):
- \` src/__tests__/auth.test.ts:describe("login") \` - Test structure and mocking patterns
**External References** (libraries and frameworks):
- Official docs: \` https://zod.dev/?id=basic-usage \` - Zod validation syntax
**WHY Each Reference Matters** (explain the relevance):
- Don't just list files - explain what pattern/information the executor should extract
- Bad: \` src/utils.ts \` (vague, which utils? why?)
- Good: \` src/utils/validation.ts:sanitizeInput() \` - Use this sanitization pattern for user input
**Acceptance Criteria**:
2026-02-02 14:18:01 +09:00
> **AGENT-EXECUTABLE VERIFICATION ONLY** — No human action permitted.
> Every criterion MUST be verifiable by running a command or using a tool.
2026-02-01 16:47:50 +09:00
**If TDD (tests enabled):**
- [ ] Test file created: src/auth/login.test.ts
- [ ] bun test src/auth/login.test.ts → PASS (3 tests, 0 failures)
2026-02-16 15:19:31 +09:00
**QA Scenarios (MANDATORY — task is INCOMPLETE without these):**
2026-02-01 16:47:50 +09:00
2026-02-16 15:19:31 +09:00
> **This is NOT optional. A task without QA scenarios WILL BE REJECTED.**
>
> Write scenario tests that verify the ACTUAL BEHAVIOR of what you built.
> Minimum: 1 happy path + 1 failure/edge case per task.
> Each scenario = exact tool + exact steps + exact assertions + evidence path.
>
> **The executing agent MUST run these scenarios after implementation.**
> **The orchestrator WILL verify evidence files exist before marking task complete.**
2026-02-02 14:18:01 +09:00
2026-02-01 16:47:50 +09:00
\\ \` \\ \` \\ \`
2026-02-16 15:19:31 +09:00
Scenario: [Happy path — what SHOULD work]
Tool: [Playwright / interactive_bash / Bash (curl)]
Preconditions: [Exact setup state]
2026-02-02 14:18:01 +09:00
Steps:
2026-02-16 15:19:31 +09:00
1. [Exact action — specific command/selector/endpoint, no vagueness]
2. [Next action — with expected intermediate state]
3. [Assertion — exact expected value, not "verify it works"]
Expected Result: [Concrete, observable, binary pass/fail]
Failure Indicators: [What specifically would mean this failed]
Evidence: .sisyphus/evidence/task-{N}-{scenario-slug}.{ext}
Scenario: [Failure/edge case — what SHOULD fail gracefully]
Tool: [same format]
Preconditions: [Invalid input / missing dependency / error state]
2026-02-02 14:18:01 +09:00
Steps:
2026-02-16 15:19:31 +09:00
1. [Trigger the error condition]
2. [Assert error is handled correctly]
Expected Result: [Graceful failure with correct error message/code]
Evidence: .sisyphus/evidence/task-{N}-{scenario-slug}-error.{ext}
2026-02-01 16:47:50 +09:00
\\ \` \\ \` \\ \`
2026-02-16 15:19:31 +09:00
> **Anti-patterns (your scenario is INVALID if it looks like this):**
> - ❌ "Verify it works correctly" — HOW? What does "correctly" mean?
> - ❌ "Check the API returns data" — WHAT data? What fields? What values?
> - ❌ "Test the component renders" — WHERE? What selector? What content?
> - ❌ Any scenario without an evidence path
2026-02-01 16:47:50 +09:00
**Evidence to Capture:**
2026-02-02 14:18:01 +09:00
- [ ] Each evidence file named: task-{N}-{scenario-slug}.{ext}
2026-02-16 15:19:31 +09:00
- [ ] Screenshots for UI, terminal output for CLI, response bodies for API
2026-02-01 16:47:50 +09:00
**Commit**: YES | NO (groups with N)
- Message: \` type(scope): desc \`
- Files: \` path/to/file \`
- Pre-commit: \` test command \`
---
2026-02-16 15:19:31 +09:00
## Final Verification Wave (MANDATORY — after ALL implementation tasks)
> **ALL 4 review agents run in PARALLEL after every implementation task is complete.**
> **ALL 4 must APPROVE before the plan is considered done.**
> **If ANY agent rejects, fix issues and re-run the rejecting agent(s).**
- [ ] F1. Plan Compliance Audit
**Agent**: oracle (read-only consultation)
**What this agent does**:
Read the original work plan (.sisyphus/plans/{name}.md) and verify EVERY requirement was fulfilled.
**Exact verification steps**:
1. Read the plan file end-to-end
2. For EACH item in "Must Have": verify the implementation exists and works
- Run the verification command listed in "Definition of Done"
- Check the file/endpoint/feature actually exists (read the file, curl the endpoint)
3. For EACH item in "Must NOT Have": verify it was NOT implemented
- Search codebase for forbidden patterns (grep, ast_grep_search)
- If found → REJECT with specific file:line reference
4. For EACH TODO task: verify acceptance criteria were met
- Check evidence files exist in .sisyphus/evidence/
- Verify test results match expected outcomes
5. Compare final deliverables against "Concrete Deliverables" list
**Output format**:
\\ \` \\ \` \\ \`
## Plan Compliance Report
### Must Have: [N/N passed]
- [✅/❌] [requirement]: [evidence]
### Must NOT Have: [N/N clean]
- [✅/❌] [guardrail]: [evidence]
### Task Completion: [N/N verified]
- [✅/❌] Task N: [criteria status]
### VERDICT: APPROVE / REJECT
### Rejection Reasons (if any): [specific issues]
\\ \` \\ \` \\ \`
- [ ] F2. Code Quality Review
**Agent**: unspecified-high
**What this agent does**:
Review ALL changed/created files for production readiness. This is NOT a rubber stamp.
**Exact verification steps**:
1. Run full type check: \` bunx tsc --noEmit \` (or project equivalent) → must exit 0
2. Run linter if configured: \` bunx biome check . \` / \` bunx eslint . \` → must pass
3. Run full test suite: \` bun test \` → all tests pass, zero failures
4. For EACH new/modified file, check:
- No \` as any \` , \` @ts-ignore \` , \` @ts-expect-error \`
- No empty catch blocks \` catch(e) {} \`
- No console.log left in production code (unless intentional logging)
- No commented-out code blocks
- No TODO/FIXME/HACK comments without linked issue
- Consistent naming with existing codebase conventions
- Imports are clean (no unused imports)
5. Check for AI slop patterns:
- Excessive inline comments explaining obvious code
- Over-abstraction (unnecessary wrapper functions)
- Generic variable names (data, result, item, temp)
**Output format**:
\\ \` \\ \` \\ \`
## Code Quality Report
### Build: [PASS/FAIL] — tsc exit code, error count
### Lint: [PASS/FAIL] — linter output summary
### Tests: [PASS/FAIL] — N passed, N failed, N skipped
### File Review: [N files reviewed]
- [file]: [issues found or "clean"]
### AI Slop Check: [N issues]
- [file:line]: [pattern detected]
### VERDICT: APPROVE / REJECT
\\ \` \\ \` \\ \`
- [ ] F3. Real Manual QA
**Agent**: unspecified-high (with \` playwright \` skill if UI involved)
**What this agent does**:
Actually RUN the deliverable end-to-end as a real user would. No mocks, no shortcuts.
**Exact verification steps**:
1. Start the application/service from scratch (clean state)
2. Execute EVERY QA scenario from EVERY task in the plan sequentially:
- Follow the exact steps written in each task's QA Scenarios section
- Capture evidence (screenshots, terminal output, response bodies)
- Compare actual behavior against expected results
3. Test cross-task integration:
- Does feature A work correctly WITH feature B? (not just in isolation)
- Does the full user flow work end-to-end?
4. Test edge cases not covered by individual tasks:
- Empty state / first-time use
- Rapid repeated actions
- Invalid/malformed input
- Network interruption (if applicable)
5. Save ALL evidence to .sisyphus/evidence/final-qa/
**Output format**:
\\ \` \\ \` \\ \`
## Manual QA Report
### Scenarios Executed: [N/N passed]
- [✅/❌] Task N - Scenario name: [result]
### Integration Tests: [N/N passed]
- [✅/❌] [flow name]: [result]
### Edge Cases: [N tested]
- [✅/❌] [case]: [result]
### Evidence: .sisyphus/evidence/final-qa/
### VERDICT: APPROVE / REJECT
\\ \` \\ \` \\ \`
- [ ] F4. Scope Fidelity Check
**Agent**: deep
**What this agent does**:
Verify that EACH task implemented EXACTLY what was specified — no more, no less.
Catches scope creep, missing features, and unauthorized additions.
**Exact verification steps**:
1. For EACH completed task in the plan:
a. Read the task's "What to do" section
b. Read the actual diff/files created for that task (git log, git diff, file reads)
c. Verify 1:1 correspondence:
- Everything in "What to do" was implemented → no missing features
- Nothing BEYOND "What to do" was implemented → no scope creep
d. Read the task's "Must NOT do" section
e. Verify NONE of the forbidden items were implemented
2. Check for unauthorized cross-task contamination:
- Did Task 5 accidentally implement something that belongs to Task 8?
- Are there files modified that don't belong to any task?
3. Verify each task's boundaries are respected:
- No task touches files outside its stated scope
- No task implements functionality assigned to a different task
**Output format**:
\\ \` \\ \` \\ \`
## Scope Fidelity Report
### Task-by-Task Audit: [N/N compliant]
- [✅/❌] Task N: [compliance status]
- Implemented: [list of what was done]
- Missing: [anything from "What to do" not found]
- Excess: [anything done that wasn't in "What to do"]
- "Must NOT do" violations: [list or "none"]
### Cross-Task Contamination: [CLEAN / N issues]
### Unaccounted Changes: [CLEAN / N files]
### VERDICT: APPROVE / REJECT
\\ \` \\ \` \\ \`
---
2026-02-01 16:47:50 +09:00
## Commit Strategy
| After Task | Message | Files | Verification |
|------------|---------|-------|--------------|
| 1 | \` type(scope): desc \` | file.ts | npm test |
---
## Success Criteria
### Verification Commands
\` \` \` bash
command # Expected: output
\` \` \`
### Final Checklist
- [ ] All "Must Have" present
- [ ] All "Must NOT Have" absent
- [ ] All tests pass
\` \` \`
---
`