- New src/agents/sisyphus/kimi-k2-6.ts based on gpt-5-4.ts 8-block architecture
- New src/agents/sisyphus-junior/kimi-k2-6.ts worker variant
- Preserves 4 pillars: intent gate + verbalization, parallel tools, verification
- Adds <re_entry_rule>: suppress re-verbalization for already-resolved turns
- Adds <exploration_budget>: hard stop conditions alongside aggressive parallelism
- Tiered <verification_loop> (V1/V2/V3): V3 keeps full rigor with harsh enforcement
- Adds <token_economy>: verbalization explicitly excluded from trim mandate
- isKimiK2Model in types.ts: matches kimi, k2p5/k2p6 variants (case-insensitive)
- Routing in sisyphus.ts + sisyphus-junior/agent.ts
- Tests: 3 new kimi routing cases in sisyphus-junior/index.test.ts (all pass)
Motivation: K2.x was post-trained with Toggle RL (~25-30% token reduction) and a
GRM scoring appropriate detail + intent inference. Reusing Claude-style prompts
double-taxes the model — external strictness on top of RL-learned strictness causes
over-deliberation on already-resolved requests. The re-entry rule and exploration
budget fix this without weakening verification rigor.
Refs: kimi.com/blog/kimi-k2-6, arxiv 2602.02276 §4.4.2 (Toggle, GRM)
The Codex 5.2 restyle in ad9df3f68 watered down the four deep-work
exhortations from gpt-5.4 (tool_call_philosophy, tool_persistence,
dependency_checks, dig_deeper) into a single bullet, leaving the
'deep worker' identity without behavioral teeth.
Restore them as Codex-style sub-sections under Exploration:
- Tool-call discipline: more calls = more accuracy, retry on partial,
read more files than needed.
- Dig deeper: don't stop at first plausible answer, check second-order
issues, prefer root over symptom (with concrete example).
- Dependency checks: resolve prerequisites before acting.
- Anti-duplication: extracted from inline paragraph to its own block.
LSP clean. 267 -> 315 lines.
Previous prose-dense rewrite went too far in stripping bullet structure.
Codex 5.1/5.2 prompts (the closest reference for an OpenAI deep-worker
prompt) actually use bullets liberally - just well-grouped (4-6 per list)
with prose introductions on each section. Restructure 5.5 to mirror that
style and tone while preserving Hephaestus's identity and all behavioral
rules from the prior round.
Sections lifted directly from Codex 5.1/5.2 organization:
- # How you work / ## Personality at the top for tonal priming
- # AGENTS.md spec as a standalone section with its own bullets
- ## Autonomy and Persistence with prose intro + Three-attempt sub-protocol
- ## Responsiveness with Frequency, Tone, Content, Examples sub-blocks
(examples rewritten to Hephaestus voice: 'Walking the agents/ tree',
'Found the dispatch in createSisyphusAgent', etc.)
- ## Plan tool with 'use a plan when' bullet list
- ## Validating your work with approval-mode granularity
(non-interactive / interactive / test-related)
- ## Presenting your work with categorical Final answer rules
(Section Headers / Bullets / Monospace / File references / Tone /
Verbosity / Don't)
- # Tool Guidelines as separate top-level section
Hephaestus-specific content preserved verbatim:
- Forge god identity, deep-worker / executor framing
- task() restricted to research subagents only
- Three-attempt failure protocol
- End-to-end usage gate (interactive_bash / playwright / curl / driver)
- Anti-duplication rule on parallel exploration
Amp-derived rules kept compact in their own ## Pragmatism and Scope:
- Smallest correct change, duplication > premature abstraction
- Default-no-tests with explicit exceptions
- WIP-not-legacy rule
- Multi-agent dirty worktree safety
Metrics: 110 -> 267 lines (still -14% from original 312), 4 -> 100 bullets
(grouped Codex-style, not scattered), 24 headers. 38/38 verification
checks pass; LSP clean.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Prior 5.5 prompt enumerated rules across 60+ bullets and 312 lines, which
fights GPT-5.5's strength: it follows prose instructions reliably and does
not need rule-by-rule cataloging. Rewrite as flowing paragraphs while
preserving the deep-worker identity and every load-bearing behavior.
Identity preserved:
- Forge god mythology ("Your boulder is code", "forge it until done")
- Direct executor, not orchestrator (research subagents only)
- Senior-colleague tone, end-to-end persistence
Behaviors preserved (compressed to prose):
- Three-attempt failure protocol → 1 paragraph
- Anti-duplication on parallel exploration
- End-to-end usage gate (interactive_bash / playwright / curl / driver)
- Implementation gate: when delegated, execute directly, no draft loop
Net additions distilled from Amp + Codex 5.2 evolution:
- Pragmatism block: smallest correct change, duplication > premature
abstraction, do not over-engineer, do not validate impossible scenarios
- Default-no-tests: add tests only when user asks, fixes a subtle bug,
or protects an important boundary; never to codebases without tests
- WIP-not-legacy: earlier unreleased shapes in the same turn are drafts,
not legacy contracts requiring backward compatibility
- Multi-agent worktree: continue task without reverting unknown changes
- Code-review mode trigger: "review" → findings-first, severity-sorted
- Personality-first opener (Codex 5.2 pattern) for tonal priming
Metrics: 312 → 110 lines (-65%), 60+ bullets → 4 bullets, 21,803 → 15,654
chars (-28%). 26/26 verification checks pass.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Hephaestus is the autonomous deep-worker - everything it receives is a
delegation. Existing 'Manual behavior' bullet vaguely said 'actually run
it' but left the validation surface unspecified, which lets a checked-in
diff plus passing tests masquerade as completion on user-visible work.
Add a dedicated 'End-to-end usage is the gate' subsection in Codex prose
style (no threats/CAPS, contract frames). Surface determines tool:
- TUI / CLI → interactive_bash (tmux), drive it like a real user
- Web / browser / UI → playwright skill, drive a real browser session
- HTTP API / service → curl or integration script against running service
- Library / SDK → minimal driver script
Reinforce in Forbidden stops trailer: when receiving a delegation,
execute directly and validate through the gate; do not loop back with
a draft when the work is yours to do.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
The existing FULL DELEGATION manual-QA rule said 'use it yourself' but
left the choice of tool implicit. Make it explicit and non-optional, so
the agent cannot satisfy the gate by reading the source instead of
running the artifact.
Surface → tool mapping:
- TUI / CLI work → interactive_bash (tmux). Launch in real terminal,
send keystrokes, run happy path, try bad input, hit --help.
- Web / browser / UI work → playwright skill. Drive a real browser,
click elements, fill forms, watch console, screenshot if helpful.
- HTTP API / service work → curl or integration script against the
running service.
- Library / SDK work → minimal driver script that imports + executes.
- Other surfaces → ask how a real user would discover it works, then
do that.
Frame the gate as a contract violation when bypassed: reporting
'implementation complete' without using the matching tool is the same
failure pattern as deleting a failing test for a green build.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Existing <verification> required tests pass + lsp clean + build green, but
that is insufficient for end-to-end delegation. Tests cover known cases;
they do not cover whether the user-visible feature actually works.
Add a NON-NEGOTIABLE rule: when the user hands off end-to-end ("ulw",
"implement and finish", "do the whole thing", "make it work", "ship it"),
verification escalates to:
1. BUILD the actual artifact
2. USE IT YOURSELF as a real user would
3. VERIFY end-to-end behavior matches the spec
4. TASK NOT DONE until usage confirms it works
Reporting "implementation complete" without having USED the artifact is
explicitly framed as a contract violation. Defects discovered during this
QA pass are the agent's to fix in the same turn.
This complements the existing 'lsp_diagnostics catches type errors, not
logic bugs' line by giving full-delegation cases a sharper, named gate.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Commit 708891dab fixed most test expectations after the gpt-5.5 model
promotion but missed 13 tests across 6 files that still expected
openai/gpt-5.4 in DEFAULT_CATEGORIES and AGENT_MODEL_REQUIREMENTS.
Updates all remaining stale expectations to openai/gpt-5.5:
- agents/utils.test.ts: atlas/metis resolution, buildAgent category,
override.category expansion (5 tests)
- plugin-handlers/config-handler.test.ts: ultrabrain config resolution
and fallback (2 tests)
- shared/agent-variant.test.ts: sisyphus chain variant and category
fallback (2 tests)
- shared/model-capability-guardrails.test.ts: built-in requirement
model ID assertion (1 test)
- tools/look-at/multimodal-fallback-chain.test.ts: multimodal-looker
hardcoded variant metadata (1 test)
- cli/config-manager/generate-omo-config.test.ts: sisyphus model and
fallback_models expectations (2 tests)
Updates test expectations across agent, cli, shared, plugin, and tools tests
to match gpt-5.5 as the new default for oracle, hephaestus, and deep agents.
Includes snapshot updates for model-fallback tests.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Inline ORACLE_GPT_5_5_PROMPT constant added to oracle.ts (Oracle is
a single-file agent, no sub-directory variant split).
Distinctive elements over ORACLE_GPT_PROMPT:
- Confidence signaling (high/medium/low) added as a required field
alongside the existing effort estimate (borrowed from Codex's
review_prompt.md)
- Codex-style section headers (# General, ## Decision framework,
## Response structure, etc.) replacing the XML-tagged structure
- Prose-first output more explicitly encouraged
- Three-tier response structure (Essential / Expanded / Edge cases)
preserved with the same hard numerical limits
- Follow-up session behavior explicitly documented
createOracleAgent() branches on isGpt5_5Model first, then isGptModel,
falling back to the thinking-enabled claude default.
The base prompt is category-agnostic; the actual category context (deep,
quick, ultrabrain, writing) layers on top at runtime via the
promptAppend parameter resolved by resolvePromptAppend.
Distinctive elements:
- Closing '# Category context' section explicitly telling the agent
to read the appended block as overriding defaults on conflict
- Orchestrator-facing final-answer structure (What changed / Key
decisions / Verification / Observations / Blockers) instead of a
user-facing conversational close
- Sparse commentary cadence; the orchestrator synthesizes progress
for the user, so mid-task narration is mostly noise
getSisyphusJuniorPromptSource() checks gpt-5-5 before the gpt-5.4 /
gpt-5.3-codex path so the new prompt takes precedence for gpt-5.5
deployments.
Ground-up rewrite that follows the same Codex-style section structure
as the new gpt-5-5 sisyphus prompt, tuned for Hephaestus's autonomous
deep-worker role.
Distinctive elements:
- 'Autonomy and Persistence' section with named 'Forbidden stops' list
(replaces gpt-5-4's FORBIDDEN/CORRECT table rhetoric)
- 'Three-attempt failure protocol' codified
- 'Exploration-first approach' with explicit 5-15 minute expectation
- 'Dig deeper' subsection for root-cause bias
- 'Ambition vs precision' distinction for greenfield vs existing
codebase work
getHephaestusPromptSource() now checks gpt-5-5 before gpt-5-4; the
regex-based gpt-5-4 path stays as the catch-all for other native
versions.
Ground-up rewrite styled after OpenAI Codex's gpt-5.4 prompt
architecture: '# General' -> '## Autonomy and Persistence' -> '## Task
execution' -> '## Validating your work' -> '# Working with the user' ->
'# Tool Guidelines' section hierarchy.
Key differences from the gpt-5-4 variant:
- Prose-first output, bullets only when content is list-shaped
- Contract frames replace threat frames (GPT-5.5 follows instructions
well; NEVER/FORBIDDEN rhetoric adds entropy without compliance gain)
- Explicit opener blacklist for 'Done -', 'Got it', 'Great question'
- '{{ personality }}' slot reserved for future persona substitution
- '{{ taskSystemGuide }}' slot switches todo/task tools per harness cfg
- Codex-compatible clickable file reference format
Sisyphus factory now checks isGpt5_5Model before isGptNativeSisyphusModel,
so gpt-5.5 models route to the new prompt while gpt-5.4, gpt-5.6+, and
other matches stay on the existing gpt-5-4 prompt.
The GPT_NATIVE_SISYPHUS_RE regex already matches gpt-5.5 (and future
5.6+), which is correct for shared behavior. However, gpt-5.5 now has
its own prompt family separate from gpt-5.4, so we need a narrower
check to route exclusively to the gpt-5-5 variants before falling
through to the regex-matched gpt-5-4 path.
The regex stays as the catch-all for future versions; isGpt5_5Model
is the precise-match guard for the current release.
Replace isGpt5_4Model + isGpt5_5Model + OR-composed isGptNativeSisyphusModel
with a single regex matching GPT-5.x where x >= 4. Automatically covers
future versions (5.6, 5.7, 5.10+) without code changes.
Constraint: Must continue to reject gpt-5.3-codex and gpt-5.x where x < 4
Rejected: Per-version functions | not scalable, each new version adds a function + OR clause
Confidence: high
Scope-risk: narrow
ast_grep_search was mentioned exactly once in the librarian prompt
(bundled as 'grep/ast_grep_search for function/class') with no syntax
guidance. When the librarian cloned a repo and tried to match code
shape, it fell into the same regex-in-AST trap as the main agent.
Two targeted edits, no rewrite of the surrounding request-classification
flow:
- Phase 1 TYPE B 'Find the implementation' now separates ast_grep_search
(code shape) from grep (text/literals) and reminds the LLM that AST
patterns use $VAR and $$$ and are not regex.
- TOOL REFERENCE adds a dedicated ast_grep_search row with valid
examples and the explicit regex anti-pattern list, alongside tightened
guidance for grep_app and grep so the LLM picks the right tool for
cross-repo vs single-repo, text vs shape.
The previous Tool Strategy was a neutral 5-bullet list that treated
ast_grep_search and grep as equals. LLMs read 'structural patterns
(function shapes, class structures)' and reach for ast_grep_search
first, then call it with regex ('foo|bar', '.*', '\\w') and silently
get zero results.
Rewrite so the default is clear - grep first, ast_grep_search only for
true AST shape matching - and enumerate the regex anti-patterns with
their corrective switches. Add an explicit rule: if ast_grep_search
returns zero matches and the printed hint says the pattern is regex-
shaped, switch to grep instead of retrying with another regex variant.
Preserves the existing absolute-path requirement, <results> block
format, and read-only / no-emoji constraints.
Document the new primary chain and install-time fallback behavior for explorer and librarian.\nKeep the user-facing guidance aligned with the runtime and CLI model selection.
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Syncs the README translations, CONTRIBUTING, docs/reference,
docs/guide, docs/examples JSONC configs, and the hierarchical
src/**/AGENTS.md files with the model version bump already landed
in the source and migration commits.
Updates the canonical Anthropic Opus model in every fallback chain
(sisyphus, oracle, prometheus, metis, momus, visual-engineering,
ultrabrain, deep, artistry, unspecified-high), the unspecified-high
category default, the think-mode HIGH_VARIANT_MAP, the Claude Code
alias map, the claude-thinking legacy alias, the context-limit GA
regex, and event.ts fallback strings.
Widens supportsCachedAnthropicLimit to accept both claude-*-4-6 and
claude-*-4-7 so the 1M context cache still applies across the bump.
Regenerates the bundled model-capabilities snapshot from models.dev
and the model-fallback snapshot to match the new source output.
Several places still emitted task(session_id=...) after the refactor:
- src/hooks/atlas/verification-reminders.ts: 2 occurrences
- src/agents/dynamic-agent-core-sections.ts: buildNonClaudePlannerSection prompt
Tests updated to match: atlas index.test.ts and dynamic-agent-prompt-builder.test.ts
The agent prompt described PDF handling but did not tell the agent
to call the Read tool, which is its only allowed tool. Added
explicit instruction so PDFs and documents are actually loaded
before extraction.
🤖 Generated with OhMyOpenCode assistance
https://github.com/code-yeongyu/oh-my-opencode
Anthropic API accepts claude-opus-4.6 (dot format) but not
claude-opus-4-6 (dash format). Added anthropic case to
transformModelForProvider to normalize dash-format model IDs before
they reach the provider. Updated model config snapshots and test
expectations to match the new dot-format output.
🤖 Generated with OhMyOpenCode assistance
https://github.com/code-yeongyu/oh-my-opencode