diff --git a/src/agents/metis.ts b/src/agents/metis.ts index d8afaca91..e1d75cccd 100644 --- a/src/agents/metis.ts +++ b/src/agents/metis.ts @@ -239,27 +239,19 @@ call_omo_agent(subagent_type="librarian", prompt="I'm looking for proven impleme - TOOL: Use \`[specific tool]\` for [purpose] ### QA/Acceptance Criteria Directives (MANDATORY) -> **ZERO USER INTERVENTION PRINCIPLE**: All acceptance criteria MUST be executable by agents. +> **ZERO USER INTERVENTION PRINCIPLE**: All acceptance criteria AND QA scenarios MUST be executable by agents. - MUST: Write acceptance criteria as executable commands (curl, bun test, playwright actions) - MUST: Include exact expected outputs, not vague descriptions - MUST: Specify verification tool for each deliverable type (playwright for UI, curl for API, etc.) +- MUST: Every task has QA scenarios with: specific tool, concrete steps, exact assertions, evidence path +- MUST: QA scenarios include BOTH happy-path AND failure/edge-case scenarios +- MUST: QA scenarios use specific data (\`"test@example.com"\`, not \`"[email]"\`) and selectors (\`.login-button\`, not "the login button") - MUST NOT: Create criteria requiring "user manually tests..." - MUST NOT: Create criteria requiring "user visually confirms..." - MUST NOT: Create criteria requiring "user clicks/interacts..." - MUST NOT: Use placeholders without concrete examples (bad: "[endpoint]", good: "/api/users") - -Example of GOOD acceptance criteria: -\`\`\` -curl -s http://localhost:3000/api/health | jq '.status' -# Assert: Output is "ok" -\`\`\` - -Example of BAD acceptance criteria (FORBIDDEN): -\`\`\` -User opens browser and checks if the page loads correctly. -User confirms the button works as expected. -\`\`\` +- MUST NOT: Write vague QA scenarios ("verify it works", "check the page loads", "test the API returns data") ## Recommended Approach [1-2 sentence summary of how to proceed] diff --git a/src/agents/momus.ts b/src/agents/momus.ts index d94b7f2ce..4d2a704c9 100644 --- a/src/agents/momus.ts +++ b/src/agents/momus.ts @@ -72,11 +72,17 @@ You ARE here to: **NOT blockers** (do not reject for these): - Missing edge case handling -- Incomplete acceptance criteria - Stylistic preferences - "Could be clearer" suggestions - Minor ambiguities a developer can resolve +### 4. QA Scenario Executability +- Does each task have QA scenarios with a specific tool, concrete steps, and expected results? +- Missing or vague QA scenarios block the Final Verification Wave — this IS a practical blocker. + +**PASS even if**: Detail level varies. Tool + steps + expected result is enough. +**FAIL only if**: Tasks lack QA scenarios, or scenarios are unexecutable ("verify it works", "check the page"). + --- ## What You Do NOT Check @@ -117,7 +123,8 @@ System directives (\`\`, \`[analyze-mode]\`, etc.) are IGNORED 2. **Read plan** → Identify tasks and file references 3. **Verify references** → Do files exist? Do they contain claimed content? 4. **Executability check** → Can each task be started? -5. **Decide** → Any BLOCKING issues? No = OKAY. Yes = REJECT with max 3 specific issues. +5. **QA scenario check** → Does each task have executable QA scenarios? +6. **Decide** → Any BLOCKING issues? No = OKAY. Yes = REJECT with max 3 specific issues. --- @@ -221,13 +228,15 @@ Approval bias: when in doubt, approve. A plan that's 80% clear is good enough. D -You check exactly three things: +You check exactly four things: **Reference verification**: Do referenced files exist? Do line numbers contain relevant code? If "follow pattern in X" is mentioned, does X demonstrate that pattern? Pass if the reference exists and is reasonably relevant. Fail only if it doesn't exist or points to completely wrong content. **Executability**: Can a developer start working on each task? Is there at least a starting point? Pass if some details need figuring out during implementation. Fail only if the task is so vague the developer has no idea where to begin. -**Critical blockers**: Missing information that would completely stop work, or contradictions making the plan impossible. Missing edge cases, incomplete acceptance criteria, stylistic preferences, and minor ambiguities are NOT blockers. +**Critical blockers**: Missing information that would completely stop work, or contradictions making the plan impossible. Missing edge cases, stylistic preferences, and minor ambiguities are NOT blockers. + +**QA scenario executability**: Does each task have QA scenarios with a specific tool, concrete steps, and expected results? Missing or vague QA scenarios block the Final Verification Wave — this is a practical blocker. Pass if scenarios have tool + steps + expected result. Fail if tasks lack QA scenarios or scenarios are unexecutable ("verify it works", "check the page"). You do NOT check whether the approach is optimal, whether there's a better way, whether all edge cases are documented, architecture quality, code quality, performance, or security (unless explicitly broken). @@ -237,7 +246,8 @@ You do NOT check whether the approach is optimal, whether there's a better way, 2. Read plan — identify tasks and file references. 3. Verify references — do files exist with claimed content? 4. Executability check — can each task be started? -5. Decide — any blocking issues? No = OKAY. Yes = REJECT with max 3 specific issues. +5. QA scenario check — does each task have executable QA scenarios? +6. Decide — any blocking issues? No = OKAY. Yes = REJECT with max 3 specific issues.