When the first prompt fails before any durable user message persists,
runtime fallback retry was rebuilding the request from parts alone and
losing the delegated agent system prompt and tool gates. Now it threads
bootstrap.system and bootstrap.tools into the retry body alongside the
captured retry parts, so the retried prompt keeps the same scope as the
initial delegate launch.
Add optional system and tools fields to DelegatedChildSessionBootstrap
so callers can stash the original delegated context alongside retry
parts. Backward compatible - existing callers stay unchanged.
- script/run-ci-tests.ts: CI test sharding and isolation logic
- script/run-ci-tests.test.ts: tests for CI test target selection
- src/features/background-agent/session-route.ts: session prompt routing for background agents
- src/hooks/interactive-bash-session/parser.ts: interactive bash output parser
- src/hooks/ralph-loop/completion-promise-detector-test-input.ts: test fixture for completion promise detection
Adds a new `notepad-write-guard` hook that intercepts Write tool calls
whose target path matches `**/.sisyphus/notepads/**` and throws an
actionable error instead of allowing the write to proceed.
Without this guard, an agent that hits an Edit hash-mismatch failure
could silently fall back to Write, destroying the entire history of an
append-only notepad file (decisions.md, issues.md, etc.). The file
carries an explicit "NEVER overwrite" warning that the agent ignores
under context pressure.
The guard is path-based so it works regardless of plan name or nesting
depth. Non-notepad `.sisyphus/**` paths (e.g. plan files) are
unaffected. The hook is wired into `create-tool-guard-hooks` under the
hook name `notepad-write-guard` and follows the same safeCreateHook +
HookName schema pattern as every other tool-guard hook.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add optional `displayName` field to AgentOverrideConfigSchema (next to
the existing `color` field) so users can specify localized agent names
in oh-my-openagent.json:
{ "agents": { "sisyphus": { "displayName": "总指挥" } } }
When set, the override takes precedence everywhere
AGENT_DISPLAY_NAMES[agentName] is used — TUI agent selector, the
agent list key, and the internal `name` field. When not set, behavior
is identical to before (hardcoded English names from AGENT_DISPLAY_NAMES).
Implementation touches:
- AgentOverrideConfigSchema: adds displayName?: z.string().optional()
- getAgentDisplayName / getAgentListDisplayName: accept optional overrides
map and check displayName before the hardcoded table
- remapAgentKeysToDisplayNames: forwards overrides map to name resolution
- agent-config-handler: passes pluginConfig.agents as the overrides map
Backward compatible — existing configs without displayName continue to
work unchanged.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
oh-my-openagent injects MCP servers (websearch, context7, grep_app) at
runtime via the OpenCode plugin API. The `opencode mcp list` command reads
only OpenCode's static config and therefore reports no servers even though
the plugin MCPs are active — this is expected, not a bug.
Add a "Native vs plugin-injected MCPs" subsection to docs/reference/features.md
that explains the three-tier architecture, shows the visibility table, and
points users to `bunx oh-my-openagent doctor --verbose` for runtime
inspection. Add brief inline notes in README.md at both MCP bullet points.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Two P0 fixes for the assistant loop that repeats its final summary 3-5 times
after all todos are marked completed before stagnation detection finally halts it.
P0.1 — session-level stop flag: when handleSessionIdle detects incompleteCount===0
it now sets state.allTodosCompletedAt. Subsequent idle events for the same session
bail out immediately at the top of the function before any HTTP fetch or injection
logic runs, preventing the re-entry loop regardless of todo-fetch caching latency.
The flag is cleared by resetContinuationProgress so sessions that receive new todos
after completion resume enforcement normally.
P0.2 — snapshot comparison scope: getTodoSnapshot now only serialises the
{id → status} mapping (sorted by key). Content and priority changes are excluded
from the comparison. Previously those fields were included, causing hasTodoSnapshotChanged
to return true whenever the LLM re-wrote todo text with identical status — which
reported progressSource="todo" and reset stagnationCount to 0, preventing
MAX_STAGNATION_COUNT=3 from ever being reached.
P1 fixes (CONTINUATION_PROMPT adversarial wording, 10 s completion grace period)
are deferred to a follow-up PR as noted in the issue.
Preserve delegated child prompt/bootstrap metadata for early runtime fallback before OpenCode has persisted the first user turn. Bind prompt gate calls to the SDK session receiver and keep completed background task lookup visible across plugin manager instances.
When a team-mode subagent hit a fallback model (rate limit / quota
exhaustion on the primary), the fallback continuation started a
fresh subagent session that was not registered in the team's
member registry under the original role. Subsequent
team_send_message / team_status calls from the fallback agent
threw "not in team" because the membership lookup missed.
Capture teamRunId + member identity at fallback initiation and
carry them onto the fallback session so the fallback agent
remains a first-class team participant. If preservation is not
possible, surface a bounded structured error instead of letting
the runtime fail mid-flight with a confusing membership message.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
When oh-my-opencode created a delegated tmux pane, terminal
capability/color probe replies emitted by tmux or the freshly
attaching opencode session could end up in the caller pane's
input buffer instead of being consumed by the delegated pane,
appearing as literal text in the main OpenCode chat (e.g.
"414/21212a2/...").
Root cause: buildSplitArgs in team-layout-tmux/layout.ts called
split-window without the -d (detached/don't-switch-focus) flag.
Without -d, tmux briefly grants focus to the new pane during
creation; the outer terminal then sends DA1/DA2 and OSC color
probe replies into what it believes is the active pane, but the
focus handoff races and those bytes land in the caller pane's
stdin buffer instead.
Fix: add -d to every split-window call in buildSplitArgs, matching
the same flag already used in pane-spawn.ts for inline subagent
panes. This keeps the caller pane focused throughout the delegated
pane lifecycle so probe replies are consumed by the correct target.
Existing tests pass; one new test asserts -d is present on every
split-window call to guard this invariant.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When a hard-reject agent (e.g. prometheus) called team_create with an
explicit `lead` in the spec, the eligibility check that runs in the
no-lead branch of shouldReuseCallerLeadSession was bypassed. The
caller session was never registered in the team, the spawned lead
ran as a detached child, and replies routed to the spawned lead
never reached the caller — the caller became an orphan that could
send but never receive.
Move the caller eligibility guard to the top of team_create.execute
so it runs unconditionally before any team-run state mutates. Throw
an actionable error naming the agent and explaining hard-reject
agents cannot lead teams regardless of an explicit `lead` in the
spec.
When a team member task errored, the failure stayed in the member's
internal state and the main/coordinator agent's wait/status loop
kept polling indefinitely — the run stalled with no visible error.
Emit a structured `member_error` peer_message into the team mailbox
on member error transition, naming the member and including the
underlying error text so the main agent's next team_status (or
pending-message read) returns a terminal failure instead of an
empty in-progress poll.
Regression test asserts the failure is visible in the main agent's
view after a member task errors mid-execution.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When ctx.serverUrl had a port string of "0", TmuxSessionManager
silently replaced it with the localhost:4096 fallback and
createTeamLayout subsequently skipped pane creation without any
user-visible signal. The two-step silent failure made team_mode
tmux_visualization look broken in default TUI mode.
Surface the failure path:
- TmuxSessionManager now retains ctx.serverUrl on the instance and
exposes it via getCtxServerUrl(), and emits a structured warning
log on the port-0 fallback branch naming both the discarded URL
and the fallback it landed on.
- createTeamLayout's "opencode server not reachable" log is
upgraded to a structured warning including ctxServerUrl and a
hint to launch with --port N + OPENCODE_PORT=N.
No behavior change to the fallback resolution itself - only the
silence. Existing port-0 fallback tests still pass; two new tests
assert the warning fires on port 0 and is absent for real ports.
Add optional enabled_expansions field to keyword_detector config schema.
When set, acts as an allowlist - only those expansion types fire.
Empty array disables all expansions. Absent field keeps all enabled (backward-compatible).
Also supports coexistence with disabled_keywords denylist for fine-grained control.
- src/config/schema/keyword-detector.ts: add enabled_expansions field
- src/hooks/keyword-detector/detector.ts: apply allowlist filter in detectKeywordsWithType
- src/hooks/keyword-detector/hook.ts: pass enabled_expansions from config
- src/hooks/keyword-detector/index.test.ts: add 4 tests for enabled_expansions behavior
- assets/oh-my-opencode.schema.json: regenerate schema
Documents the v4.2.0 release window in Keep-a-Changelog format, including prompt gate fixes, internal audits, known issues, and the watchdog supersession history.
Closes LOW-14, LOW-16
New AST-based audit walks all *.test.ts files under src/ and asserts every mock.module(...) call is paired with cleanup. Existing offenders are documented in MOCK_MODULE_LIFECYCLE_ALLOWLIST with TODO references.
Closes HIGH-10
Race-condition and concurrency fixes must include reporter-verified repro confirmation before the originating issue is closed. Adds the checklist and rationale grounded in recent incident examples.
Closes MEDIUM-12
Documents the reservation-based duplicate-injection guard introduced in v4.2.0 with accepted status, exported API signatures, release semantics, migration notes, and commit references.
Closes MEDIUM-11
The promptWithModelSuggestionRetry async variant did not release the
post-dispatch reservation when the wrapped promptAsync threw. Callers
that immediately retry (such as sendSyncPrompt error toast paths) hit
the gate as reserved and surfaced 'promptAsync skipped by gate: reserved'
instead of the underlying error.
Mirrors the existing sync variant fix from ff1b15d53.
Closes regression introduced by BLOCKER-2 hardening
Lines 79/142/428 of prompt-async-gate.test.ts used timer-based synchronization, violating .sisyphus/rules/test-discipline.md which forbids time-based test waits. Replace them with explicit dispatch awaits and mocked-time expiry so the assertions do not depend on CI machine speeds.
Closes BLOCKER-3 (Wave 2 cleanup)
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Replace the inlined parent-wake coalescing logic in manager.ts with delegation to the ParentWakeNotifier extracted in c1ccf8d09. The four timer Maps and the related methods now live in their own module with a narrow public API, while BackgroundManager retains the wiring point and the enqueue-callback bridge.
Closes HIGH-9 (step 2: integration)
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
PR #3825 introduced a delegated child-session bootstrap to capture first-prompt retry payloads before history is persisted, addressing the empty-history fallback gap. After merge the PR's own regression test failed on clean root bun test (6828 pass / 1 fail), so PR #4044 reverted it. Ship v4.2.0 with the bug documented and a workaround so users have an explicit story for the unfixed delegated child-session early-failure path. Reland will target v4.2.1.
Closes BLOCKER-4 (Path B - reland deferred to v4.2.1)
Document all 7+ BLOCKER+HIGH fixes, breaking-change-free additions
(public exports), known issue for delegated child-session fallback
(PR #3825 deferred to v4.2.1), and internal-only changes.
Closes L14
Walk all test files, parse with TypeScript Compiler API, assert every
mock.module(path, factory) invocation has a paired afterEach/afterAll
cleanup. Existing offenders are allowlisted with TODOs for v4.2.1 work.
Closes H10
Test-discipline.md forbids setTimeout(resolve, N) and sleep(N) in test bodies. Replace the 3 microtask and expiry sleeps with explicit microtask yields and deterministic clock advancement, preserving the prompt gate invariants without real-time waits.
Closes BLOCKER-3
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
PR #3825 added a shared bootstrap context to capture delegated
child-session retry payloads before the first prompt dispatch, so
empty-history failures could still retry through the fallback chain.
The PR's own regression test failed on clean root bun test after merge
(6828 pass / 1 fail). PR #4044 reverted the merge to keep dev green.
Ship v4.2.0 with the bug documented and a workaround so users have an
explicit story for the unfixed delegated child-session early-failure
path. Reland targets v4.2.1 once the regression test is stabilized.
Closes BLOCKER-4 (Path B - documentation, reland deferred to v4.2.1)