Source code change: - src/shared/model-requirements.ts: prepend claude-sonnet-4-6 to metis fallback chain so Sonnet becomes the default. Opus 4.7 max remains as the immediate fallback for callers who want extra reasoning. - src/shared/model-requirements.test.ts: update assertion to expect Sonnet primary + Opus secondary. AGENTS.md accuracy fixes (verified against source): - Agent modes: Sisyphus/Hephaestus are 'primary' (not 'all'); Sisyphus-Junior is 'subagent' (not 'all'). Confirmed via 'const MODE: AgentMode = ...' in each agent file. Also clarified Prometheus has no agentSources factory and is built via buildPrometheusAgentConfig. - Sisyphus fallback chain: corrected order to kimi-k2.6 → k2p5 → kimi-k2.5 → gpt-5.5 medium → glm-5 → big-pickle (was missing kimi-k2.5). - Librarian/Explore: added missing minimax-m2.7 step between -highspeed and claude-haiku-4-5. - Metis chain: removed fictitious gemini-3.1-pro entry. - Sisyphus-Junior chain: spelled out the actual fallback (was 'user-configurable'). - Temperatures: Sisyphus/Hephaestus do not set explicit temperature (model default); Sisyphus-Junior is 0.1 via SISYPHUS_JUNIOR_DEFAULTS. - Quick category default: gpt-5.4-mini (not gpt-5.4-mini-fast). Team-mode corrections: - Eligibility registry has 3 verdicts: eligible (sisyphus, atlas, sisyphus-junior), conditional (hephaestus — needs D-36 teammate permission), hard-reject (oracle, librarian, explore, multimodal-looker, metis, momus, prometheus). - Schema has 11 fields, not 4: added max_messages_per_run, max_wall_clock_minutes, max_member_turns, base_dir, message_payload_max_bytes, recipient_unread_max_bytes, mailbox_poll_interval_ms. - Hooks: 'team-session-events' is 4 sub-handlers in src/plugin/event.ts (team-idle-wake-hint, team-lead-orphan-handler, team-member-error-handler, team-member-status-handler), not a single Continuation-tier hook. - Tier counts now show base + team-mode: ToolGuard 14/15, Transform 5/7. - Total: 52 base hooks, 59 with team-mode. Doc cascade for the Metis change: - docs/guide/orchestration.md, agent-model-matching.md, installation.md - docs/reference/configuration.md, features.md
6.5 KiB
GPT-5.5 System Prompt Drafts
This directory contains ground-up rewrites of the Sisyphus, Hephaestus, Oracle, and Deep system prompts, styled after OpenAI Codex's gpt-5.4 prompt architecture and targeted at GPT-5.5.
Files
sisyphus.md— Orchestrator. Intent gate, delegation philosophy, parallel execution discipline, verification.hephaestus.md— Autonomous deep worker. Persistence, exploration-first, forbidden stops, root-cause bias.oracle.md— Read-only strategic advisor. Three-tier response structure, hard verbosity limits, confidence signaling.deep.md— Category-spawned deep worker (runs as Sisyphus-Junior under thedeepcategory). Goal-oriented autonomous execution.
Design principles applied
Each prompt applies the same small set of principles, borrowed and adapted from Codex's gpt-5.4 prompt work:
- Single identity header with
{{ personality }}slot. Separates persona from logic so the same base prompt can ship in default / friendly / pragmatic variants without duplication. # General→## Autonomy and Persistence→## Task execution→## Validating your work→# Working with the user→# Tool Guidelinesstructure. Lifted directly from Codex'sgpt_5_2_prompt.mdandgpt-5.2-codex_prompt.md. Keeps the same section contract for every agent so readers can navigate consistently.- Prose-first output, bullets only when list-shaped. GPT-5.5 reads and writes prose naturally; bullet overuse is a GPT-5.3 coping mechanism, not a genuine formatting need.
- Contract frames over threat frames. Rules are stated as agreements and expectations, not as "NEVER DO X OR YOU WILL FAIL". GPT-5.5's instruction following is strong enough that threats add entropy without improving compliance.
- Opener blacklist is explicit. "Done —", "Got it", "Great question", "Sure thing", and similar filler are called out by name. These are the most common failure modes across all models.
- File reference formatting is unified. Clickable markdown links with absolute paths, no
file://orhttps://for local files, no line ranges. - Why, not just what. Each major rule is accompanied by the reasoning. Rules without reasons get ignored when models judge them weakly-grounded; rules with reasons get applied even in novel situations.
Agent-specific shape
Sisyphus
- Intent classification table (surface form → true intent → routing).
- Zero-tolerance visual-engineering delegation rule.
- Six-section delegation prompt contract.
- Session continuity (
task_idreuse) as a first-class topic. - Oracle consultation as a separate section with clear use/not-use guidance.
Hephaestus
- Forbidden stops as a named list.
- Three-attempt failure protocol.
- Exploration-first as explicit philosophy (5-15 minutes is normal).
- "Dig deeper" subsection for root-cause bias.
- Ambition vs precision distinction for greenfield vs existing codebase work.
- Task-tool restriction stated as an intentional design decision with rationale.
Oracle
- Three-tier response structure (Essential / Expanded / Edge cases) with hard numerical limits.
- Effort estimation (Quick / Short / Medium / Large) as a required field.
- Confidence signaling (high / medium / low) added as a required field — new in v5.5, borrowed from Codex's
review_prompt.md. - Pragmatic minimalism as explicit decision framework.
- "No commentary channel; every word is the final answer" constraint acknowledged.
Deep
- Explicitly positioned as Sisyphus-Junior in
deepmode (category-spawned counterpart to Hephaestus). - Extensive exploration expectation stated.
- Final-answer structure tuned for orchestrator relay: "What changed / Key decisions / Verification / Observations / Blockers".
- Commentary cadence tuned down (sparse) since the user is not directly on the other side.
Known deviations from Codex
These are intentional choices where oh-my-opencode's architecture differs from Codex's:
task()delegation is central for Sisyphus (it is the orchestrator), entirely absent for Oracle (read-only consultant), research-only for Hephaestus and Deep (they execute directly).- No
update_plantool; the harness usestask_create/task_updateinstead. Each prompt references its own tool set. - Sub-agent ecosystem (explore, librarian, oracle, metis, momus) is specific to this harness and does not exist in Codex. Each prompt explains when and how to use these agents.
- Skill loading is a first-class concept via the
skilltool. Codex has a simpler skill model. - Commentary / final channels are named the same way as Codex's output contract, but the actual transport layer is different (OpenCode, not Codex CLI).
Line counts
For reference, approximate line counts after this rewrite versus the current production prompts:
| Agent | Current (assembled) | Draft | Delta |
|---|---|---|---|
| Sisyphus GPT-5.4 | ~500 | ~270 | -46% |
| Hephaestus GPT-5.4 | ~400 | ~270 | -33% |
| Oracle GPT | ~120 | ~160 | +33% |
| Deep category append | ~20 | ~250 (as standalone) | N/A |
Oracle grew because v5.5 adds Confidence signaling and explicitly documents follow-up session behavior. Deep grew because the draft is a standalone prompt rather than a category append; in production it would either replace Sisyphus-Junior's GPT-5.5 variant entirely or layer on top of a minimal Sisyphus-Junior base.
What this draft is not
- Not a
.tsfile. These are markdown drafts. Converting to TypeScript template strings (with{todoHookNote},{keyTriggers}, etc. interpolation) is the next step, once the content is validated. - Not a tested prompt. These have not been run against evals. Before shipping, each prompt should be benchmarked with
skill-creator's eval loop against the current production prompts on a representative task set. - Not personality-substituted. The
{{ personality }}slot is a placeholder. Default / friendly / pragmatic content still needs to be authored.
Suggested next steps
- Author personality variants. Three short paragraphs (default, friendly, pragmatic) that slot into
{{ personality }}and can be reused across all four prompts. - Build an eval harness. Pick 5-10 representative tasks per agent and run current-prod vs draft-v5.5 head-to-head.
- Convert to
.tswith dynamic composition helpers. Preserve the existingbuildAgentIdentitySection,buildToolSelectionTable, etc. integration points where they still apply. - Ship behind a feature flag. Opt-in for
gpt-5.5model selection until eval confidence is high.