Commit Graph

1658 Commits

Author SHA1 Message Date
ZeyuFu a8ccffdd7c style(runtime-fallback): add explicit optional chain on .replace per review
Addresses cubic-dev-ai P1 finding on #4113 (#4113 review).

The original chain `extractErrorName(error)?.toLowerCase().replace(...)`
is semantically safe — JavaScript optional chaining short-circuits the
ENTIRE access chain when the head returns null/undefined, so when
`extractErrorName` returns undefined the whole expression evaluates to
undefined without ever reaching `.replace()`. Verified empirically via
`const x = undefined; x?.toLowerCase().replace(/_/g, "")` returns
undefined with no crash.

Applying the suggested defensive `?.` before `.replace` anyway, since
it is semantically a no-op and explicit chaining at each hop is easier
for static analyzers to reason about.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-17 11:05:01 -04:00
ZeyuFu 2bd7946001 fix(non-interactive-env): use powershell syntax on Windows regardless of SHELL/MSYSTEM
Closes #3607.

## Root cause

On Windows, `detectCommandShellType()` fell through to `detectShellType()`
for two common environments and incorrectly returned `"unix"`:

1. **SHELL points at a Unix-shaped path** (e.g. Git Bash sets
   `SHELL=/usr/bin/bash` on a fresh Windows install).
   The `detectWindowsShellType(process.env.SHELL)` probe didn't recognize
   `bash` as a Windows shell, so the function fell through and
   `detectShellType()` returned `"unix"`.

2. **MSYSTEM is set but SHELL is not** (Git Bash leaves MSYSTEM permanently
   set system-wide even when the active shell is PowerShell).
   The fall-through path returned `"unix"` via the MSYSTEM check.

In both cases, the hook then prepended `export KEY=val;` to git commands,
which PowerShell rejects with:

  `export : 无法将"export"项识别为 cmdlet...`

OpenCode on Windows runs the bash tool through a Windows shell
(PowerShell by default, cmd as the user-overridable fallback), regardless
of MSYSTEM or a Unix-shaped SHELL set by Git Bash — so the env prefix
must use Windows-compatible syntax.

## Fix

`detectCommandShellType()` now short-circuits on `process.platform === "win32"`:

- If `SHELL` points at a recognized Windows shell (`cmd.exe`, `powershell.exe`,
  `pwsh.exe`), return that.
- If `SHELL` and `MSYSTEM` are both unset, fall back to `ComSpec` then to cmd.
- Otherwise, default to PowerShell — matching what OpenCode actually spawns.

`detectShellType()` is unchanged; other callers (including non-Windows
platforms) are unaffected.

## Test changes

Three pre-existing tests encoded the buggy behavior as expected behavior
and have been updated to assert the new PowerShell syntax with a
`(#3607)` marker and a comment explaining why a Unix-shaped SHELL on
win32 must still resolve to PowerShell. WSL is not affected because in
WSL `process.platform === "linux"`, not `"win32"`.

- `src/hooks/non-interactive-env/`: 24 tests pass / 0 fail
- `bunx tsc --noEmit`: clean

## Note on issue thread

The sisyphus-bot triage comment on #3607 framed this as a policy choice
between (A) forcing Windows env-prefix syntax and (B) resolving against
the OpenCode-configured shell. This PR implements option (A) as the
minimal surgical fix; option (B) remains a follow-up.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-17 10:59:29 -04:00
ZeyuFu b2f0d42394 test(runtime-fallback): tighten quota regression fixtures so new paths actually fire
Addresses cubic-dev-ai bot review on #4113 (P2): the RESOURCE_EXHAUSTED
and snake_case insufficient_quota fixtures contained quota-shaped
messages that already matched pre-existing message regexes, so the tests
passed even without the new errorName allow-list entry and the
underscore normalization respectively.

Replace both fixture messages with a generic "Request failed." so the
only path to a `quota_exceeded` classification is via the new code:

- RESOURCE_EXHAUSTED: only the new `errorName?.includes("resourceexhausted")`
  match on the normalized name can fire.
- insufficient_quota (snake_case): only the new underscore-stripping
  normalization can route the name to `insufficientquota` and match the
  existing allow-list entry.

The third new test (Google ResourceExhausted message-only) is unchanged
because its message uniquely matches only the new
`/resource.?exhausted/i` pattern and not any existing quota regex.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-17 10:37:40 -04:00
ZeyuFu f357ed033a fix(runtime-fallback): classify more provider quota error names
Closes #3937.

Adds three small classification gaps to `classifyErrorType` so that
quota-exhaustion errors from a wider range of providers trigger
configured fallback chains instead of looping retry attempts:

- Normalize error names by stripping `_` and `-` so snake_case /
  SCREAMING_SNAKE_CASE provider names (`insufficient_quota`,
  `RESOURCE_EXHAUSTED`, `rate_limit_exceeded`) match the existing
  alphanumeric `.includes()` checks.
- Add `resourceexhausted` to the quota error-name allow-list to cover
  Google Generative AI's gRPC code 8 / `ResourceExhausted` surface.
- Add `/resource.?exhausted/i` to the quota message-pattern list so the
  same error surface is caught when the provider only sets a generic
  error name but puts the signal in the message.

Three new regression tests in
`quota-error-classifier.regression.test.ts` cover:

- Google `RESOURCE_EXHAUSTED` (gRPC error name + quota-shaped message)
- Google `ResourceExhausted` message form without HTTP status
- OpenAI snake_case `insufficient_quota` error name

No existing tests were touched; the underscore normalization preserves
all existing `.includes()` matches by rewriting the one underscore-bearing
literal (`ai_loadapikeyerror` → `ailoadapikeyerror`) so previously
matched names still resolve.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-17 10:28:05 -04:00
ZeyuFu 6c54123ec1 fix(slash-commands): inject command content exactly once (#3724)
Guard command.execute.before against injecting when parts already
contain auto-slash-command tags, preventing duplication when both
chat.message and command.execute.before fire for the same slash command.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-17 07:07:49 -04:00
deopa0402 9758168676 test(auto-update): isolate cached version resolution
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-05-17 18:36:42 +09:00
deopa0402 37d9d613b6 fix(auto-update): clean stale OMO cache roots
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-05-17 18:35:29 +09:00
YeonGyu-Kim 6768decddb fix(session-recovery): fallback when stored unavailable-tool parts are absent
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-05-17 17:16:10 +09:00
YeonGyu-Kim 12bd658079 refactor(prompt-async-gate): remove deprecated dispatch wrappers
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-05-17 17:15:54 +09:00
YeonGyu-Kim 1bbe065c60 refactor(prompt-callers): migrate shared and cli dispatch
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-05-17 17:09:04 +09:00
YeonGyu-Kim 989ab7171d refactor(hooks): use unified internal prompt dispatch
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-05-17 17:07:35 +09:00
YeonGyu-Kim b5d24619c8 test(prompt-async-gate): pin unified internal prompt dispatch contract
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-05-17 16:34:05 +09:00
YeonGyu-Kim 8bc4977563 fix(slash-command): skip already tagged command output 2026-05-17 16:17:42 +09:00
YeonGyu-Kim 55312cc4b6 fix(session-recovery): preflight idle recovery fanout 2026-05-17 16:17:36 +09:00
YeonGyu-Kim a7b7ace7ed fix(prompt-gate): block prompts into pending tool turns 2026-05-17 15:42:58 +09:00
YeonGyu-Kim 6eb88a0545 fix(session-recovery): prefer valid tool use ids 2026-05-17 15:15:13 +09:00
YeonGyu-Kim f43effb842 fix(session-recovery): recover interrupted idle tool turns 2026-05-17 15:08:39 +09:00
Disaster-Terminator ec371237d3 Merge remote-tracking branch 'origin/dev' into fix/task-id-prompt-surface
# Conflicts:
#	src/agents/atlas/default-prompt-sections.ts
#	src/agents/atlas/gemini-prompt-sections.ts
#	src/agents/atlas/gpt-prompt-sections.ts
#	src/agents/hephaestus/gpt-5-3-codex.ts
2026-05-17 10:17:24 +08:00
YeonGyu-Kim 75223149dd Merge pull request #4096 from code-yeongyu/kimi-k2.6
fix: add merge-conflict guard test to prevent source file corruption
2026-05-17 04:32:40 +09:00
YeonGyu-Kim d8f365bfd4 test(guard): add merge-conflict guard to prevent unresolved git conflicts in source files
Unresolved git merge conflict markers (<<<<<<<, =======, >>>>>>>) in
TypeScript source files break parsing and can cause the plugin to fail
at runtime or tests to hang with cryptic errors. This guard scans all
.ts/.tsx/.json files under src/ and fails the test suite if any
conflict markers are found.

Closes #debugging-hang-investigation
2026-05-17 04:17:45 +09:00
YeonGyu-Kim a77312c371 test(atlas): exclude isSessionActive timeout from retry timer assertion
The status timeout uses setTimeout internally, which the test's mocked
setTimeout captures. Filter it alongside the dispatch timeout.
2026-05-17 03:56:39 +09:00
YeonGyu-Kim fcd0011a6b test(atlas): track active timers instead of scheduled delays in setTimeout mock
- Replace the simple scheduledDelays array with an activeTimers Map so
  that clearTimeout removes timers from the tracked set.
- This prevents false positives when internal withDispatchTimeout calls
  setTimeout for safety timeouts that are immediately cancelled.
- Keeps the test intent unchanged: only genuinely scheduled retries are
  counted as delayed duplicate retries.
2026-05-17 03:50:02 +09:00
YeonGyu-Kim 2613de522f fix(prompt-async-gate): timeout isSessionActive to prevent infinite hang on stale SDK status
- Wrap isSessionActive in withDispatchTimeout (capped at 5s) so a
  stuck OpenCode SDK status() call cannot block internal prompts forever.
- Catch the timeout and treat session as inactive so the prompt can
  proceed rather than hanging indefinitely.
- Add regression test: session.status that never resolves now times out
  and allows dispatch instead of hanging the test (and production).

Refs: AGENTS.md internal-message-injection safety note
2026-05-17 03:41:08 +09:00
YeonGyu-Kim 36468b118d fix(shared): add timeout to isSessionActive to prevent infinite hang
The SDK client disables fetch timeout (req.timeout = false). When the
opencode server is slow or unresponsive, client.session.status() hangs
forever, blocking the entire dispatchToHooks chain (sequential await)
and deadlocking the plugin.

Add a 5s timeout wrapper around the status call in isSessionActive().
On timeout, the catch block returns false (treat as inactive), letting
the caller proceed instead of hanging indefinitely.

The timeout parameter is exposed for testing to avoid 5s test delays.

Refs: .debugging/status-timeout-hang.md
2026-05-17 03:40:17 +09:00
YeonGyu-Kim 271878bcea perf(rules-injector): cache full candidates and memoize ancestor scans
The rule scan cache stored a path[] keyed by (projectRoot|startDir|
skipClaudeUserRules), so two issues stacked up on every tracked tool
call:

- Cache hits still ran safeRealpathSync(realpathSync) and re-derived
  isGlobal / distance / isSingleFile for every cached path. That is a
  per-candidate sync syscall plus repeated string-prefix walks.
- Sibling files in the same project landed under different startDir
  keys, so the entire walk-and-recursive-scan chain repeated even
  though every ancestor rule directory was identical.

Store the full RuleFileCandidate[] in the per-call cache so a cache hit
returns immediately with no realpath syscall. Add a separate per-
directory scan cache (getDirScan/setDirScan) keyed by absolute rule
directory path, so two sibling files reuse the same readdir + realpath
work for every shared ancestor.

Microbench (200 files / 20 modules / cached session):
- single sweep: 41.8ms -> 2.5ms (16x)
- 3-pass replay: 88.6ms -> 3.2ms (28x)

Pin the new invariants with two new tests:
- 'does not re-resolve symlinked rule path on cache hit' via a
  retargeted directory symlink.
- 'reuses ancestor directory scan for sibling files in the same
  project' by deleting the source rule file between the two calls.
2026-05-17 02:06:52 +09:00
YeonGyu-Kim c25f75294e perf(rules-injector): cache project root for visited ancestors
findProjectRoot was keyed by exact startPath, so sibling files in the
same project repeated the entire upward marker walk. The walk does one
existsSync per marker per ancestor directory, which adds up on every
read/write/edit/multiedit tool call.

Track every directory visited during the walk and seed the cache with
the resolved root for each of them. Subsequent lookups for any
descendant short-circuit to the cached ancestor without re-running
marker probes. Cache invalidation still happens on session.deleted /
session.compacted, so production semantics are unchanged.

Pin the new contract via a sibling-startpath test, and make the
existing finder.test.ts beforeEach explicit about cache state so the
more aggressive cache does not leak between tests.
2026-05-17 02:06:52 +09:00
YeonGyu-Kim 25d8054192 Merge pull request #4074 from code-yeongyu/fix/delegate-task-spawn
fix(delegate-task): start child prompts reliably
2026-05-17 01:00:03 +09:00
YeonGyu-Kim ba648685d4 fix(runtime-fallback): carry delegated system and tools through bootstrap retry
When the first prompt fails before any durable user message persists,
runtime fallback retry was rebuilding the request from parts alone and
losing the delegated agent system prompt and tool gates. Now it threads
bootstrap.system and bootstrap.tools into the retry body alongside the
captured retry parts, so the retried prompt keeps the same scope as the
initial delegate launch.
2026-05-17 00:08:30 +09:00
YeonGyu-Kim 4f9813848a feat: add ci test runner, session routing, bash parser, and test fixtures
- script/run-ci-tests.ts: CI test sharding and isolation logic
- script/run-ci-tests.test.ts: tests for CI test target selection
- src/features/background-agent/session-route.ts: session prompt routing for background agents
- src/hooks/interactive-bash-session/parser.ts: interactive bash output parser
- src/hooks/ralph-loop/completion-promise-detector-test-input.ts: test fixture for completion promise detection
2026-05-16 23:51:43 +09:00
ZeyuFu 572c3c248e fix(notepad-guard): refuse Write tool for .sisyphus/notepads files (#3685)
Adds a new `notepad-write-guard` hook that intercepts Write tool calls
whose target path matches `**/.sisyphus/notepads/**` and throws an
actionable error instead of allowing the write to proceed.

Without this guard, an agent that hits an Edit hash-mismatch failure
could silently fall back to Write, destroying the entire history of an
append-only notepad file (decisions.md, issues.md, etc.).  The file
carries an explicit "NEVER overwrite" warning that the agent ignores
under context pressure.

The guard is path-based so it works regardless of plan name or nesting
depth.  Non-notepad `.sisyphus/**` paths (e.g. plan files) are
unaffected.  The hook is wired into `create-tool-guard-hooks` under the
hook name `notepad-write-guard` and follows the same safeCreateHook +
HookName schema pattern as every other tool-guard hook.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-16 06:52:43 -04:00
ZeyuFu 57ab749b37 refactor(keyword-detector): consolidate analyze pattern/message into analyze/default and document delegate_task params (#3694)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-16 06:50:46 -04:00
ZeyuFu 5e4d45a3ca fix(todo-continuation-enforcer): stop looping after all todos complete (#4013)
Two P0 fixes for the assistant loop that repeats its final summary 3-5 times
after all todos are marked completed before stagnation detection finally halts it.

P0.1 — session-level stop flag: when handleSessionIdle detects incompleteCount===0
it now sets state.allTodosCompletedAt. Subsequent idle events for the same session
bail out immediately at the top of the function before any HTTP fetch or injection
logic runs, preventing the re-entry loop regardless of todo-fetch caching latency.
The flag is cleared by resetContinuationProgress so sessions that receive new todos
after completion resume enforcement normally.

P0.2 — snapshot comparison scope: getTodoSnapshot now only serialises the
{id → status} mapping (sorted by key). Content and priority changes are excluded
from the comparison. Previously those fields were included, causing hasTodoSnapshotChanged
to return true whenever the LLM re-wrote todo text with identical status — which
reported progressSource="todo" and reset stagnationCount to 0, preventing
MAX_STAGNATION_COUNT=3 from ever being reached.

P1 fixes (CONTINUATION_PROMPT adversarial wording, 10 s completion grace period)
are deferred to a follow-up PR as noted in the issue.
2026-05-16 06:31:18 -04:00
YeonGyu-Kim 5a2c3bbba8 fix(workspace): match omo guard paths cross-platform 2026-05-16 19:14:54 +09:00
YeonGyu-Kim 240a4a17ad fix(workspace): harden omo migration review issues 2026-05-16 18:02:54 +09:00
YeonGyu-Kim 82ec099c3a fix(atlas): match omo as a path segment 2026-05-16 17:52:28 +09:00
YeonGyu-Kim 63519ec563 docs(workspace): document omo workspace paths 2026-05-16 17:42:06 +09:00
YeonGyu-Kim f10f796318 fix(workspace): keep omo and legacy rules compatible 2026-05-16 17:41:49 +09:00
YeonGyu-Kim 36e373cdbb feat(workspace): point planning guardrails at omo 2026-05-16 17:41:35 +09:00
YeonGyu-Kim a86221b1a9 feat(workspace): store runtime state under omo 2026-05-16 17:40:11 +09:00
YeonGyu-Kim 982fa81367 fix(delegate-task): start child prompts reliably
Preserve delegated child prompt/bootstrap metadata for early runtime fallback before OpenCode has persisted the first user turn. Bind prompt gate calls to the SDK session receiver and keep completed background task lookup visible across plugin manager instances.
2026-05-16 16:27:56 +09:00
YeonGyu-Kim d974cd3d3b test(hooks): repair stale retry harnesses 2026-05-16 15:43:15 +09:00
YeonGyu-Kim cf7bf9d02d fix(team-mode): accept legacy inline specs 2026-05-16 15:43:08 +09:00
YeonGyu-Kim a20540579e Merge pull request #4068 from code-yeongyu/feat/pre-publish-fix-v420
v4.2.0: pre-publish review fixes (BLOCKER-1..3, HIGH-5..10, MID-11/12)
2026-05-16 14:50:54 +09:00
ZeyuFu 2bf5038215 fix(team-mode): surface member error to main agent (#3923)
When a team member task errored, the failure stayed in the member's
internal state and the main/coordinator agent's wait/status loop
kept polling indefinitely — the run stalled with no visible error.
Emit a structured `member_error` peer_message into the team mailbox
on member error transition, naming the member and including the
underlying error text so the main agent's next team_status (or
pending-message read) returns a terminal failure instead of an
empty in-progress poll.

Regression test asserts the failure is visible in the main agent's
view after a member task errors mid-execution.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-16 01:40:08 -04:00
pizzav-xyz f152569ed4 feat(keyword-detector): add enabled_expansions config for allowlist control
Add optional enabled_expansions field to keyword_detector config schema.
When set, acts as an allowlist - only those expansion types fire.
Empty array disables all expansions. Absent field keeps all enabled (backward-compatible).
Also supports coexistence with disabled_keywords denylist for fine-grained control.

- src/config/schema/keyword-detector.ts: add enabled_expansions field
- src/hooks/keyword-detector/detector.ts: apply allowlist filter in detectKeywordsWithType
- src/hooks/keyword-detector/hook.ts: pass enabled_expansions from config
- src/hooks/keyword-detector/index.test.ts: add 4 tests for enabled_expansions behavior
- assets/oh-my-opencode.schema.json: regenerate schema
2026-05-16 02:01:29 +02:00
YeonGyu-Kim 5a8bd05db0 test(prompt-async-gate): replace timer waits with deterministic sync (BLOCKER-3)
Lines 79/142/428 of prompt-async-gate.test.ts used timer-based synchronization, violating .sisyphus/rules/test-discipline.md which forbids time-based test waits. Replace them with explicit dispatch awaits and mocked-time expiry so the assertions do not depend on CI machine speeds.

Closes BLOCKER-3 (Wave 2 cleanup)

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-05-16 01:50:11 +09:00
YeonGyu-Kim 845d862b9b test(prompt-async-gate): replace setTimeout sleeps with deterministic sync
Test-discipline.md forbids setTimeout(resolve, N) and sleep(N) in test bodies. Replace the 3 microtask and expiry sleeps with explicit microtask yields and deterministic clock advancement, preserving the prompt gate invariants without real-time waits.

Closes BLOCKER-3

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-05-16 01:36:57 +09:00
YeonGyu-Kim f93d7297c8 test(prompt-async-gate): cover dispatch timeout and post-dispatch error hold
Adds regression coverage for BLOCKER-1 (dispatch timeout releases
reservation for next caller after stalled upstream) and BLOCKER-2
(post-dispatch error preserves the post-dispatch hold so an immediate
second caller observes the reservation and is gated).

Both tests subscribe-first on the promptAsync call count and assert
status transitions without sleep-based synchronization. dispatchTimeoutMs
is the system under test, so passing it explicitly as 1ms in those tests
is the SUT, not a sleep-as-synchronization (per test-discipline.md).

Closes BLOCKER-3 (dispatch timeout + post-dispatch coverage)

Co-authored-by: gate-tests (deep / gpt-5.3-codex high)
2026-05-16 00:49:28 +09:00
YeonGyu-Kim be25109f6d fix(continuation): skip internal user turns 2026-05-15 23:16:23 +09:00
YeonGyu-Kim c580b8f2ce fix(session): ignore internal synthetic turns 2026-05-15 23:16:05 +09:00