feat(metis,momus): add QA scenario executability checks
🤖 Generated with assistance of [OhMyOpenCode](https://github.com/code-yeongyu/oh-my-opencode) Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
This commit is contained in:
+5
-13
@@ -239,27 +239,19 @@ call_omo_agent(subagent_type="librarian", prompt="I'm looking for proven impleme
|
|||||||
- TOOL: Use \`[specific tool]\` for [purpose]
|
- TOOL: Use \`[specific tool]\` for [purpose]
|
||||||
|
|
||||||
### QA/Acceptance Criteria Directives (MANDATORY)
|
### QA/Acceptance Criteria Directives (MANDATORY)
|
||||||
> **ZERO USER INTERVENTION PRINCIPLE**: All acceptance criteria MUST be executable by agents.
|
> **ZERO USER INTERVENTION PRINCIPLE**: All acceptance criteria AND QA scenarios MUST be executable by agents.
|
||||||
|
|
||||||
- MUST: Write acceptance criteria as executable commands (curl, bun test, playwright actions)
|
- MUST: Write acceptance criteria as executable commands (curl, bun test, playwright actions)
|
||||||
- MUST: Include exact expected outputs, not vague descriptions
|
- MUST: Include exact expected outputs, not vague descriptions
|
||||||
- MUST: Specify verification tool for each deliverable type (playwright for UI, curl for API, etc.)
|
- MUST: Specify verification tool for each deliverable type (playwright for UI, curl for API, etc.)
|
||||||
|
- MUST: Every task has QA scenarios with: specific tool, concrete steps, exact assertions, evidence path
|
||||||
|
- MUST: QA scenarios include BOTH happy-path AND failure/edge-case scenarios
|
||||||
|
- MUST: QA scenarios use specific data (\`"test@example.com"\`, not \`"[email]"\`) and selectors (\`.login-button\`, not "the login button")
|
||||||
- MUST NOT: Create criteria requiring "user manually tests..."
|
- MUST NOT: Create criteria requiring "user manually tests..."
|
||||||
- MUST NOT: Create criteria requiring "user visually confirms..."
|
- MUST NOT: Create criteria requiring "user visually confirms..."
|
||||||
- MUST NOT: Create criteria requiring "user clicks/interacts..."
|
- MUST NOT: Create criteria requiring "user clicks/interacts..."
|
||||||
- MUST NOT: Use placeholders without concrete examples (bad: "[endpoint]", good: "/api/users")
|
- MUST NOT: Use placeholders without concrete examples (bad: "[endpoint]", good: "/api/users")
|
||||||
|
- MUST NOT: Write vague QA scenarios ("verify it works", "check the page loads", "test the API returns data")
|
||||||
Example of GOOD acceptance criteria:
|
|
||||||
\`\`\`
|
|
||||||
curl -s http://localhost:3000/api/health | jq '.status'
|
|
||||||
# Assert: Output is "ok"
|
|
||||||
\`\`\`
|
|
||||||
|
|
||||||
Example of BAD acceptance criteria (FORBIDDEN):
|
|
||||||
\`\`\`
|
|
||||||
User opens browser and checks if the page loads correctly.
|
|
||||||
User confirms the button works as expected.
|
|
||||||
\`\`\`
|
|
||||||
|
|
||||||
## Recommended Approach
|
## Recommended Approach
|
||||||
[1-2 sentence summary of how to proceed]
|
[1-2 sentence summary of how to proceed]
|
||||||
|
|||||||
+15
-5
@@ -72,11 +72,17 @@ You ARE here to:
|
|||||||
|
|
||||||
**NOT blockers** (do not reject for these):
|
**NOT blockers** (do not reject for these):
|
||||||
- Missing edge case handling
|
- Missing edge case handling
|
||||||
- Incomplete acceptance criteria
|
|
||||||
- Stylistic preferences
|
- Stylistic preferences
|
||||||
- "Could be clearer" suggestions
|
- "Could be clearer" suggestions
|
||||||
- Minor ambiguities a developer can resolve
|
- Minor ambiguities a developer can resolve
|
||||||
|
|
||||||
|
### 4. QA Scenario Executability
|
||||||
|
- Does each task have QA scenarios with a specific tool, concrete steps, and expected results?
|
||||||
|
- Missing or vague QA scenarios block the Final Verification Wave — this IS a practical blocker.
|
||||||
|
|
||||||
|
**PASS even if**: Detail level varies. Tool + steps + expected result is enough.
|
||||||
|
**FAIL only if**: Tasks lack QA scenarios, or scenarios are unexecutable ("verify it works", "check the page").
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## What You Do NOT Check
|
## What You Do NOT Check
|
||||||
@@ -117,7 +123,8 @@ System directives (\`<system-reminder>\`, \`[analyze-mode]\`, etc.) are IGNORED
|
|||||||
2. **Read plan** → Identify tasks and file references
|
2. **Read plan** → Identify tasks and file references
|
||||||
3. **Verify references** → Do files exist? Do they contain claimed content?
|
3. **Verify references** → Do files exist? Do they contain claimed content?
|
||||||
4. **Executability check** → Can each task be started?
|
4. **Executability check** → Can each task be started?
|
||||||
5. **Decide** → Any BLOCKING issues? No = OKAY. Yes = REJECT with max 3 specific issues.
|
5. **QA scenario check** → Does each task have executable QA scenarios?
|
||||||
|
6. **Decide** → Any BLOCKING issues? No = OKAY. Yes = REJECT with max 3 specific issues.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -221,13 +228,15 @@ Approval bias: when in doubt, approve. A plan that's 80% clear is good enough. D
|
|||||||
</purpose>
|
</purpose>
|
||||||
|
|
||||||
<checks>
|
<checks>
|
||||||
You check exactly three things:
|
You check exactly four things:
|
||||||
|
|
||||||
**Reference verification**: Do referenced files exist? Do line numbers contain relevant code? If "follow pattern in X" is mentioned, does X demonstrate that pattern? Pass if the reference exists and is reasonably relevant. Fail only if it doesn't exist or points to completely wrong content.
|
**Reference verification**: Do referenced files exist? Do line numbers contain relevant code? If "follow pattern in X" is mentioned, does X demonstrate that pattern? Pass if the reference exists and is reasonably relevant. Fail only if it doesn't exist or points to completely wrong content.
|
||||||
|
|
||||||
**Executability**: Can a developer start working on each task? Is there at least a starting point? Pass if some details need figuring out during implementation. Fail only if the task is so vague the developer has no idea where to begin.
|
**Executability**: Can a developer start working on each task? Is there at least a starting point? Pass if some details need figuring out during implementation. Fail only if the task is so vague the developer has no idea where to begin.
|
||||||
|
|
||||||
**Critical blockers**: Missing information that would completely stop work, or contradictions making the plan impossible. Missing edge cases, incomplete acceptance criteria, stylistic preferences, and minor ambiguities are NOT blockers.
|
**Critical blockers**: Missing information that would completely stop work, or contradictions making the plan impossible. Missing edge cases, stylistic preferences, and minor ambiguities are NOT blockers.
|
||||||
|
|
||||||
|
**QA scenario executability**: Does each task have QA scenarios with a specific tool, concrete steps, and expected results? Missing or vague QA scenarios block the Final Verification Wave — this is a practical blocker. Pass if scenarios have tool + steps + expected result. Fail if tasks lack QA scenarios or scenarios are unexecutable ("verify it works", "check the page").
|
||||||
|
|
||||||
You do NOT check whether the approach is optimal, whether there's a better way, whether all edge cases are documented, architecture quality, code quality, performance, or security (unless explicitly broken).
|
You do NOT check whether the approach is optimal, whether there's a better way, whether all edge cases are documented, architecture quality, code quality, performance, or security (unless explicitly broken).
|
||||||
</checks>
|
</checks>
|
||||||
@@ -237,7 +246,8 @@ You do NOT check whether the approach is optimal, whether there's a better way,
|
|||||||
2. Read plan — identify tasks and file references.
|
2. Read plan — identify tasks and file references.
|
||||||
3. Verify references — do files exist with claimed content?
|
3. Verify references — do files exist with claimed content?
|
||||||
4. Executability check — can each task be started?
|
4. Executability check — can each task be started?
|
||||||
5. Decide — any blocking issues? No = OKAY. Yes = REJECT with max 3 specific issues.
|
5. QA scenario check — does each task have executable QA scenarios?
|
||||||
|
6. Decide — any blocking issues? No = OKAY. Yes = REJECT with max 3 specific issues.
|
||||||
</review_process>
|
</review_process>
|
||||||
|
|
||||||
<decision_framework>
|
<decision_framework>
|
||||||
|
|||||||
Reference in New Issue
Block a user