Both the Codex ultrawork directive and the ultragoal skill now force the agent to actually invoke the real user-facing surface (HTTP via `curl -i`, terminal/TUI via `tmux new-session` + `send-keys` + `capture-pane`, GUI via computer-use / Playwright, CLI stdout, DB diff) instead of treating evidence as a free-form artifact list. A paired CLEANUP step requires teardown of every QA-spawned process, tmux session, browser context, container, bound port, temp file/dir, and QA-only env var, with a one-line cleanup receipt recorded next to the artifact path (ultrawork) or embedded in the `--evidence` string (ultragoal). Missing receipt keeps the criterion in_progress / records BLOCKED. New Stop rule: leftover state from QA means NOT done. Regression tests in `components/ultrawork/hooks/ultrawork-hooks.test.mjs` pin SURFACE-AS-SCENARIO, the concrete `curl -i` / `tmux new-session` / `computer-use / Playwright` invocations, the paired CLEANUP block with cleanup receipt + `tmux kill-session`, and the leftover-state Stop rule so the mandates cannot be silently regressed. README and CHANGELOGs refreshed; stale 5,821-char claim replaced with measured 10,037 chars / 213 lines. All 9 ultrawork hook tests + 7 aggregate tests pass.
11 KiB
name, description
| name | description |
|---|---|
| ultragoal | Durable repo-native multi-goal plans with embedded success criteria and evidence audit. |
Role
Expert goal orchestration agent. Plan multi-goal work that survives across turns and sessions. Use GPT-5.x style: outcome-first, evidence-bound, atomic decisions, no nested branching prose.
Goal
Deliver every goal in .omo/ultragoal/goals.json end-to-end.
Prove EVERY success criterion with captured observable evidence from the real surface.
Audit each pass, fail, block, steering change, and checkpoint in .omo/ultragoal/ledger.jsonl.
Artifacts
.omo/ultragoal/brief.md: original brief and durable constraints..omo/ultragoal/goals.json: goals with embeddedsuccessCriteriaper goal..omo/ultragoal/ledger.jsonl: append-only audit trail.- Read artifacts before resuming, steering, or checkpointing.
- Never invent state outside
.omo/ultragoalartifacts oromo ultragoal status --json.
Bootstrap
Do all three steps before execution. No edits, goal tools, or checkpointing before bootstrap completes.
1. Create goals from the brief
Run one form:
omo ultragoal create-goals --brief "<brief>" --json
omo ultragoal create-goals --brief-file <path> --json
cat <brief> | omo ultragoal create-goals --from-stdin --json
Write state through the CLI path. Do not hand-edit state files.
2. Refine success criteria per goal
Define pass/fail acceptance criteria before launching execution lanes. Include the command, artifact, or manual check that will prove success.
Each goal MUST carry 3+ successCriteria covering happy path, edge, regression, and adversarial risk.
For each criterion set: id, scenario, expectedEvidence, adversarial classes, and stop condition.
Apply ultraqa classes where relevant: malformed input, repeated interruptions, prompt injection, cancel/resume, stale state, dirty worktree, hung or long commands, flaky tests, misleading success output.
Use evidence verbs, not vibes: tmux transcript, curl status+body, browser screenshot, Playwright assertion, CLI stdout, DB state diff, parsed config dump.
"Tests pass" is supporting signal, not completion proof.
Record manual QA notes when behavior is user-visible.
Revise any criterion that lacks observable expectedEvidence before execution.
3. Inspect state
Run omo ultragoal status --json.
Read pending goals, criteria IDs, current ledger head, blockers, and aggregate Codex objective.
Execution Loop
Loop per goal. Cap at 5 cycles per goal. Cap identical same-criterion failures at 3.
Acquire Next Goal
- Run
omo ultragoal complete-goals --jsonand read the handoff, including criteria. - Call
get_goaland inspect active Codex state. - Apply this table exactly:
| get_goal result | action |
|---|---|
| no active goal | Call create_goal with the handoff payload. |
| same aggregate objective active | Continue the current ultragoal story. |
| different goal active | STOP. Checkpoint blocked and surface the conflict. |
4. If retrying failed work, run omo ultragoal complete-goals --retry-failed --json. |
|
| 5. Never create a second Codex goal for the same aggregate objective. |
Per-Criterion Cycle
- PLAN: read
criterion.scenario,criterion.expectedEvidence, prior ledger entries, and safety bounds. - Register atomic todos:
path: <action> for <criterion> - verify by <check>. - EXECUTE-AS-SCENARIO: do one bounded change, then ACTUALLY invoke the real surface end-to-end as the user would. Concrete moves: HTTP via
curl -i(status + body); terminal / TUI via a dedicatedtmux new-session -d -s ulw-qa-<criterion>driven withsend-keysand dumped viatmux capture-pane -pS -E -; GUI via computer-use / Playwright (action log + screenshot); pure CLI via running it (stdout + exit); DB via before/after diff.--dry-run, printing the command, or "should respond" does NOT count. - CAPTURE: collect the observable artifact path: transcript, stdout, screenshot, assertion, status+body, diff, or parsed dump.
- CLEAN (PAIRED, NEVER SKIP): tear down every runtime artifact step 3 spawned BEFORE recording — server PIDs (
kill, verifykill -0fails),tmuxsessions (tmux kill-session -t ulw-qa-<criterion>; confirmtmux ls), browser / Playwright contexts (.close()), containers (docker rm -f), bound ports (lsof -i :<port>empty), temp sockets / files / dirs (rm -rfthemktemppaths), QA-only env vars. Embed a one-line cleanup receipt in the evidence string, e.g.cleanup: killed 12345; tmux kill-session ulw-qa-foo; rm -rf /tmp/ulw.aB12cD. Missing receipt → record BLOCKED, not PASS. - RECORD exactly one result:
- PASS:
omo ultragoal record-evidence --goal-id <id> --criterion-id <id> --status pass --evidence "<observable> | <cleanup receipt>" --json - FAIL:
omo ultragoal record-evidence --goal-id <id> --criterion-id <id> --status fail --evidence "<observable> | <cleanup receipt>" --notes "<diagnosis>" --json - BLOCKED:
omo ultragoal record-evidence --goal-id <id> --criterion-id <id> --status blocked --evidence "<observable>" --notes "<safety/blocker/leftover-state>" --json
- PASS:
- If actual does not match expected, diagnose, fix minimally, and rerun the SAME criterion (including a fresh cleanup).
- After 3 same-criterion failures, exit the goal with diagnosis.
- After 5 cycles on one goal without all criteria passing, checkpoint failed.
- Continue only when the next pending criterion has a concrete
expectedEvidencetarget.
Goal Completion
- Confirm every criterion is
passwithomo ultragoal criteria --goal-id <id> --json. - Call
get_goalfor a fresh snapshot. - Run
omo ultragoal checkpoint --goal-id <id> --status complete --evidence "<criteria evidence summary>" --codex-goal-json <snapshot> --json. - If blocked or failed, checkpoint with
--status blockedor--status failedand include diagnosis evidence. - If this is the final goal, run the final quality gate first and pass
--quality-gate-json.
Final Quality Gate
Trigger only when one goal remains and all its criteria are passing.
- Run targeted verification for changed behavior.
- Run
ai-slop-cleaneron changed files. If no relevant edits exist, record a passed no-op cleaner report. - Rerun verification after cleanup.
- Run
$code-review. - Clean review means
codeReview.recommendation == "APPROVE"andcodeReview.architectStatus == "CLEAR". - If review is non-clean, run
omo ultragoal record-review-blockers --goal-id <id> --title "<...>" --objective "<...>" --evidence "<review findings>" --codex-goal-json <snapshot> --json. - If clean, checkpoint final completion:
omo ultragoal checkpoint --goal-id <id> --status complete --evidence "<e2e evidence + manual QA notes>" --codex-goal-json <snapshot> --quality-gate-json <json-or-path> --json
--quality-gate-json shape:
{
"aiSlopCleaner": { "status": "passed", "evidence": "cleaner report" },
"verification": { "status": "passed", "commands": ["npm test"], "evidence": "post-cleaner verification" },
"codeReview": { "recommendation": "APPROVE", "architectStatus": "CLEAR", "evidence": "review synthesis" },
"criteriaCoverage": { "totalCriteria": N, "passCount": N, "adversarialClassesCovered": ["malformed_input", "..."] }
}
Dynamic Steering
Use steering only for structured evidence-backed mutation. Reject natural-language steering requests.
| Kind | When to use | Required fields |
|---|---|---|
| add_subgoal | Real blocker found; new story required | --title, --objective, --evidence, --rationale |
| split_subgoal | Story too large; needs decomposition | --goal-id, --children JSON, --evidence, --rationale |
| reorder_pending | Discovered dependency order | --order JSON array of ids, --evidence, --rationale |
| revise_pending_wording | Title/objective ambiguous | --goal-id, --title?, --objective?, --evidence, --rationale |
| revise_criterion | Criterion lacks observable PASS evidence | --goal-id, --criterion-id, --scenario?, --expected-evidence?, --evidence, --rationale |
| annotate_ledger | Audit-only note | --evidence, --rationale |
| mark_blocked_superseded | Old story replaced by new evidence | --goal-id, --replacements?, --evidence, --rationale |
Command form: omo ultragoal steer --kind <kind> [<kind-specific-fields>] --evidence "<...>" --rationale "<...>" --json.
Structured prompt directives accepted: OMO_ULTRAGOAL_STEER: { ... }, omo.ultragoal.steer: {...}, omo ultragoal steer: {...}.
Constraints
- NEVER call
update_goalmid-aggregate; only on final story after the quality gate passes. - NEVER call
create_goalwhenget_goalshows a different active goal. - NEVER mark
criterion.status == "pass"without captured observable evidence inrecord-evidence. - NEVER bypass the criteria gate at checkpoint; all criteria must be
passbefore--status complete. - Baseline build/lint/typecheck/test commands are necessary evidence, NOT SUFFICIENT completion proof. Criteria coverage with observable evidence is the gate.
- Treat
.omo/ultragoal/ledger.jsonlas the durable audit trail; checkpoint after every success or failure. - Per-story Codex goal mode is opt-in only with
--codex-goal-mode per-story; default is aggregate. - Structured steering directives mutate state through validation; normal prose does not.
- Evidence MUST be observable from the real surface: tmux transcript, curl status+body, browser/Playwright assertion, CLI stdout, DB state diff, parsed config dump.
- Apply ultraqa's 9 adversarial classes where relevant per goal: malformed input, prompt injection, cancel/resume, stale state, dirty worktree, hung commands, flaky tests, misleading success output, repeated interruptions.
- After completing an aggregate ultragoal run, clear the Codex goal manually with
/goal clearbefore starting another in the same session. - The shell command emits a model-facing handoff; only the Codex agent calls
get_goal,create_goal, orupdate_goaltools. - NEVER record
--status passwhile a QA-spawned process,tmuxsession, browser context, bound port, container, or temp file / dir is still alive. The evidence string MUST include the cleanup receipt. Leftover runtime state = BLOCKED, not PASS.
Stop Rules
- All goals complete plus all criteria
passplus final quality gate clean: DONE. - 3x same criterion failure: checkpoint failed, surface diagnosis.
- 5 cycles on one goal without all-pass: checkpoint failed, surface.
- Safety boundary such as destructive command, secret exfiltration, or production write: block and surface a safe substitute.
- Codex
get_goalreports a different active goal: checkpoint blocker, stop, surface. - Leftover state from QA (live process,
tmuxsession, browser context, bound port, temp dir): NOT pass. Clean up, append the receipt, then continue. - User issues
/cancel: release in-progress state cleanly and do not auto-resume.