Addresses cubic-dev-ai P1 + P2 findings on #4115.
The original PR relaxed `exitCode > 1` to `> 2` based on the (wrong)
claim that ripgrep exits 2 only on non-fatal I/O issues. ripgrep
actually uses exit code 2 for BOTH fatal errors (pattern syntax,
invalid args) AND non-fatal I/O issues; GNU grep (the fallback backend
in grep/cli.ts) likewise uses 2 for fatal errors. So `> 2` would
silently suppress fatal errors.
The correct fix is just `--no-messages`, which suppresses ripgrep's
stderr only for soft I/O issues (broken symlinks, permission denied)
while leaving fatal-error messages intact. With the gate kept at
`exitCode > 1 && stderr.trim()`:
- Broken symlink: ripgrep exits 2, stderr is empty (suppressed) →
`stderr.trim()` is falsy → gate fails → partial results survive.
- Fatal error: ripgrep exits 2, stderr has the real error message
(not suppressed by --no-messages) → gate triggers → error returned.
Reverting both `exitCode > 1` → `> 2` changes; keeping the
`--no-messages` flag additions and the regression test (test comment
updated to describe the cleaner architecture).
Verification: bun test src/tools/glob/ src/tools/grep/ → 30 pass / 0 fail.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Addresses cubic-dev-ai P1 finding on #4113 (#4113 review).
The original chain `extractErrorName(error)?.toLowerCase().replace(...)`
is semantically safe — JavaScript optional chaining short-circuits the
ENTIRE access chain when the head returns null/undefined, so when
`extractErrorName` returns undefined the whole expression evaluates to
undefined without ever reaching `.replace()`. Verified empirically via
`const x = undefined; x?.toLowerCase().replace(/_/g, "")` returns
undefined with no crash.
Applying the suggested defensive `?.` before `.replace` anyway, since
it is semantically a no-op and explicit chaining at each hop is easier
for static analyzers to reason about.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes#3726.
## Root cause
Two coupled bugs in the ripgrep integration caused the glob and grep
tools to fail outright when the search path contained a broken
(dangling) symlink:
1. **Missing `--no-messages`** in `RG_FILES_FLAGS` (glob) and
`RG_SAFETY_FLAGS` (grep). ripgrep prints a stderr warning for every
broken symlink:
`rg: /path/to/broken-link: No such file or directory (os error 2)`
2. **Overly strict exit-code gate** in both `glob/cli.ts` and
`grep/cli.ts`: `if (exitCode > 1 && stderr.trim())` treated exit
code 2 as fatal and discarded the stdout, even though ripgrep exits
2 on *non-fatal* I/O issues (broken symlinks, permission denied)
while still printing valid results to stdout.
Combined effect: a single broken symlink anywhere in the search path
turned all subsequent file matches into an empty result with an error.
## Fix
- Append `--no-messages` to `RG_FILES_FLAGS` (`src/tools/glob/constants.ts`)
and to `RG_SAFETY_FLAGS` (`src/tools/grep/constants.ts`). This silences
ripgrep's non-fatal stderr warnings without changing match behavior.
- Relax the exit-code gate in both `glob/cli.ts` and `grep/cli.ts` from
`> 1` to `> 2`, matching ripgrep's documented contract:
- 0 = matches found
- 1 = no matches (success, just nothing matched)
- 2 = non-fatal I/O issues (partial success — stdout is still valid)
- >2 = fatal error
## Test changes
- New regression assertion in `src/tools/glob/cli.test.ts` for
`buildRgArgs` confirms `--no-messages` is included in the args.
- The grep change is symmetric (same flag, same rationale, same
exit-code constant) and rides on the parallel structure.
## Verification
- `bun test src/tools/glob/` → all pass (19 tests in cli.test.ts)
- `bun test src/tools/grep/` → all pre-existing tests still pass
- `bunx tsc --noEmit` → clean
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes#3607.
## Root cause
On Windows, `detectCommandShellType()` fell through to `detectShellType()`
for two common environments and incorrectly returned `"unix"`:
1. **SHELL points at a Unix-shaped path** (e.g. Git Bash sets
`SHELL=/usr/bin/bash` on a fresh Windows install).
The `detectWindowsShellType(process.env.SHELL)` probe didn't recognize
`bash` as a Windows shell, so the function fell through and
`detectShellType()` returned `"unix"`.
2. **MSYSTEM is set but SHELL is not** (Git Bash leaves MSYSTEM permanently
set system-wide even when the active shell is PowerShell).
The fall-through path returned `"unix"` via the MSYSTEM check.
In both cases, the hook then prepended `export KEY=val;` to git commands,
which PowerShell rejects with:
`export : 无法将"export"项识别为 cmdlet...`
OpenCode on Windows runs the bash tool through a Windows shell
(PowerShell by default, cmd as the user-overridable fallback), regardless
of MSYSTEM or a Unix-shaped SHELL set by Git Bash — so the env prefix
must use Windows-compatible syntax.
## Fix
`detectCommandShellType()` now short-circuits on `process.platform === "win32"`:
- If `SHELL` points at a recognized Windows shell (`cmd.exe`, `powershell.exe`,
`pwsh.exe`), return that.
- If `SHELL` and `MSYSTEM` are both unset, fall back to `ComSpec` then to cmd.
- Otherwise, default to PowerShell — matching what OpenCode actually spawns.
`detectShellType()` is unchanged; other callers (including non-Windows
platforms) are unaffected.
## Test changes
Three pre-existing tests encoded the buggy behavior as expected behavior
and have been updated to assert the new PowerShell syntax with a
`(#3607)` marker and a comment explaining why a Unix-shaped SHELL on
win32 must still resolve to PowerShell. WSL is not affected because in
WSL `process.platform === "linux"`, not `"win32"`.
- `src/hooks/non-interactive-env/`: 24 tests pass / 0 fail
- `bunx tsc --noEmit`: clean
## Note on issue thread
The sisyphus-bot triage comment on #3607 framed this as a policy choice
between (A) forcing Windows env-prefix syntax and (B) resolving against
the OpenCode-configured shell. This PR implements option (A) as the
minimal surgical fix; option (B) remains a follow-up.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Addresses cubic-dev-ai bot review on #4113 (P2): the RESOURCE_EXHAUSTED
and snake_case insufficient_quota fixtures contained quota-shaped
messages that already matched pre-existing message regexes, so the tests
passed even without the new errorName allow-list entry and the
underscore normalization respectively.
Replace both fixture messages with a generic "Request failed." so the
only path to a `quota_exceeded` classification is via the new code:
- RESOURCE_EXHAUSTED: only the new `errorName?.includes("resourceexhausted")`
match on the normalized name can fire.
- insufficient_quota (snake_case): only the new underscore-stripping
normalization can route the name to `insufficientquota` and match the
existing allow-list entry.
The third new test (Google ResourceExhausted message-only) is unchanged
because its message uniquely matches only the new
`/resource.?exhausted/i` pattern and not any existing quota regex.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes#3937.
Adds three small classification gaps to `classifyErrorType` so that
quota-exhaustion errors from a wider range of providers trigger
configured fallback chains instead of looping retry attempts:
- Normalize error names by stripping `_` and `-` so snake_case /
SCREAMING_SNAKE_CASE provider names (`insufficient_quota`,
`RESOURCE_EXHAUSTED`, `rate_limit_exceeded`) match the existing
alphanumeric `.includes()` checks.
- Add `resourceexhausted` to the quota error-name allow-list to cover
Google Generative AI's gRPC code 8 / `ResourceExhausted` surface.
- Add `/resource.?exhausted/i` to the quota message-pattern list so the
same error surface is caught when the provider only sets a generic
error name but puts the signal in the message.
Three new regression tests in
`quota-error-classifier.regression.test.ts` cover:
- Google `RESOURCE_EXHAUSTED` (gRPC error name + quota-shaped message)
- Google `ResourceExhausted` message form without HTTP status
- OpenAI snake_case `insufficient_quota` error name
No existing tests were touched; the underscore normalization preserves
all existing `.includes()` matches by rewriting the one underscore-bearing
literal (`ai_loadapikeyerror` → `ailoadapikeyerror`) so previously
matched names still resolve.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Guard command.execute.before against injecting when parts already
contain auto-slash-command tags, preventing duplication when both
chat.message and command.execute.before fire for the same slash command.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Use one dispatchInternalPrompt surface with mode: async | sync so source, settle, hold, timeout, status checks, reservations, and release semantics stay in one runner. Keep the old helper names temporarily so caller migration can land atomically in follow-up commits.
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Root cause of the user-visible `/init-deep ulw` hang (session
`ses_1cb9c3013ffesUOy5H3QOIya4K`): the plugin's `unhandledRejection` and
`uncaughtException` listeners were calling
`scheduleForcedExit(handler(error), 1, true)`, which both ran the entire
`cleanupAll()` chain (BackgroundManager shutdown, tmux pane closure,
team-mode teardown) and then `process.exit(1)`'d the host. Under heavy
slash commands like `/init-deep ultrafucking deep`, a single mid-stream
error (e.g. opencode's own `session.processor` Aborted-process condition,
or a transient socket reset) would:
1. trigger the listener,
2. abort the in-flight background tasks (`session.error
MessageAbortedError` for both child sessions in the log),
3. close the tmux panes the user was watching,
4. immediately kill the host via `process.exit(1)`.
From the user's seat that looked like a frozen TUI, which is what they
reported as "ulw 여전히 멈추는데". The error blob also logged as `{}`
because `JSON.stringify(new Error(...))` strips non-enumerable Error
fields, so the previous log line carried no diagnostic value.
This change makes the global `uncaughtException` / `unhandledRejection`
listeners log-only:
* New `describeProcessCleanupError()` extracts `{name, message, stack}`
from Error instances, falls back to a structured `{raw: ...}` payload
for plain objects / primitives, so the log now actually says what
failed.
* `registerErrorEvent()` no longer runs cleanup and no longer calls
`scheduleForcedExit`. It detaches itself, logs a single explanatory
line, and returns. Bun's default crash behaviour is already suppressed
for these events when a listener is present, so the host now genuinely
survives transient streaming errors instead of being killed by our
own helper.
* Signal handlers (`SIGINT` / `SIGTERM` / `SIGBREAK` / `beforeExit` /
`exit`) keep their existing behaviour and still run `cleanupAll()`
before the host terminates — that is now the only path that tears
down background tasks and tmux panes.
Tests are updated to lock in the new contract:
* New regression `#given scheduleForcedExit enabled AND unhandledRejection
fires #when the listener runs #then process.exit is NOT called AND
process.exitCode stays 0 AND no cleanup runs` (and the
uncaughtException twin) re-enables `scheduleForcedExit`, spies on
`process.exit` plus `globalThis.setTimeout`, and asserts none of them
are touched. Without the fix this test failed exactly like the
observed hang (exit called once, exitCode set to 1).
* The existing "manager shuts down before process exits" tests are
rewritten to assert the opposite: cleanup is NOT invoked from the
error path.
* A complementary `'exit'` listener test pins the real shutdown
contract (`exit` event still triggers `cleanupAll`).
* A new `#given describeProcessCleanupError` block covers the four
shapes (Error, plain object with own fields, empty object, primitive).
* The unregister assertion is tightened to check the listener count
drops back to baseline (previously it relied on a side effect of the
old cleanup-on-error path).
Manual QA: `bun /tmp/process-cleanup-smoke-test.ts` (out-of-tree smoke
driver) emits five back-to-back unhandledRejection/uncaughtException
events with various payload shapes and prints
`SMOKE_TEST_OK survived 5 emissions; exitCode=0; shutdownInvocations=0`,
confirming the host survives and no spurious cleanup runs.
`bun test` runs green for the affected modules:
- src/features/background-agent (528 tests)
- src/create-managers + src/plugin (218 tests)
- src/hooks/{unstable-agent-babysitter,ralph-loop,keyword-detector}
(225 tests)
Root cause: `formatTaskResult` and `formatFullSession` both call
`client.session.messages({ path: { id: task.sessionId } })` with no
timeout. The OpenCode SDK explicitly disables fetch timeout
(`req.timeout = false` in `packages/sdk/js/src/client.ts:37`), so when
`session.processor` enters an "Aborted process" loop (same condition
patched in #4086 for the Read hot path) every subsequent
`session.messages` call hangs forever.
User-visible impact: the parent agent running a heavy slash command
such as `/init-deep` fires many background `task` agents in parallel
and then calls `background_output(task_id, block=true)` for each. The
existing `timeoutMs` (max 10min) caps only the polling loop that waits
for the child task to transition out of `running`. Once the child is
`completed`, the loop exits and the code falls through to
`formatTaskResult` / `formatFullSession`. The 10min cap does not apply
to that post-completion fetch, so a wedged `session.messages` leaves
the tool call hanging indefinitely. The parent session appears stuck
to the user (the symptom they report on `/init-deep ultrafucking
deep`).
Fix: race the underlying `session.messages` call against a 5s timeout
through a new `withSdkCallTimeout` helper local to
`src/tools/background-task/` (mirrors `withFetchTimeout` in
`src/shared/dynamic-truncator.ts` and `withDispatchTimeout` in
`src/shared/prompt-async-gate.ts`). On timeout the caller returns a
clearly-marked `"Error fetching messages: ... timed out after 5000ms"`
fallback so the agent can recognise the failure and continue instead
of waiting forever.
Tests:
- New `sdk-call-timeout.test.ts` pins the fix with three BDD cases
using a never-settling mock and the new
`_setBackgroundOutputFetchTimeoutMsForTesting(50)` override:
(1) `formatTaskResult` resolves to the timeout fallback under the
fetch budget, (2) same for `formatFullSession`, (3) two parallel
callers against the wedged client both resolve cleanly.
- All three failed (timed out) before the fix and pass in 51ms each
after.
Verification:
- `bun test src/tools/background-task/` -- 33 pass.
- `bun test` (full suite) -- 7014 pass, 1 skip, 0 fail across 723
files.
- `bun run typecheck` -- clean (tsgo --noEmit).
- LSP diagnostics on the four changed files -- 0 errors.
- Manual QA harness `.local-ignore/init-deep-hang-qa.ts` (gitignored)
drove production default 5s timeout against a wedged client:
serial `formatTaskResult` resolved in 5001ms, serial
`formatFullSession` in 5002ms, 10 parallel `formatTaskResult` calls
all resolved within 5002ms with the timeout fallback. Without the
fix every call hangs indefinitely.
Refs: builds on #4086 (dynamic-truncator) which patched the Read hot
path of the same SDK hang.
Detect OpenCode promptAsync calls that return before a child session has any durable message, and surface a prompt acceptance error before the generic five-minute sync poll timeout.
Add a failing-first regression for the idle zero-message case and keep the existing durable-message completion path covered.
Debugging-Journal: .debugging
Unresolved git merge conflict markers (<<<<<<<, =======, >>>>>>>) in
TypeScript source files break parsing and can cause the plugin to fail
at runtime or tests to hang with cryptic errors. This guard scans all
.ts/.tsx/.json files under src/ and fails the test suite if any
conflict markers are found.
Closes #debugging-hang-investigation
Root cause: `getContextWindowUsage` caches the *promise* of
`fetchContextWindowUsage` in a per-session WeakMap keyed by client. When
`ctx.client.session.messages({ path: { id: sessionID } })` never settles
(observed once `service=session.processor ... error=Aborted process`
takes hold), the cached pending promise wedges every concurrent and
later caller in the same session. The five hooks that share one
`createDynamicTruncator(ctx)` -- directory-agents-injector,
directory-readme-injector, rules-injector, tool-output-truncator, plus
indirect callers -- all await that same poisoned promise on every Read,
so the user-facing tool chain hangs forever and ESC cannot break it.
Reporters in #4086 land on this path consistently when reading AGENTS.md
files (which trigger directory-agents-injector via the directory walk).
Fix: race the underlying `session.messages` call against a 5s timeout
through a new `withFetchTimeout` helper. On timeout the catch block logs
the failure and returns `null`, which `dynamicTruncate` already treats
as the "context usage unavailable" signal and falls back to the static
truncation budget. Successful responses still cache as before. The
`message.updated finish=true` invalidation hook still clears poisoned
caches on the next completed turn so retries are clean.
Tests:
- Add a never-settling `session.messages` mock with a 50 ms override via
the new `_setContextWindowUsageFetchTimeoutMsForTesting` hook (matches
the established `_setXxxForTesting` pattern in `opencode-http-api.ts`
and `prompt-async-gate.ts`).
- Three new BDD cases pin the fix: (1) single caller returns null fast,
(2) parallel concurrent callers all unblock on the same cached promise
instead of hanging, (3) invalidate + retry rehydrates cleanly.
- All 8 pre-existing tests in the file still pass (happy paths, cache
reuse, invalidation, env/model fallback).
Verification:
- `bun test src/shared/dynamic-truncator.test.ts` -- 11 pass.
- `bun test src/shared/prompt-async-route-audit.test.ts` -- 6 pass
(added log import, no raw prompt route added).
- `bun test` (full suite) -- 7009 pass, 1 skip, 1 pre-existing flake in
`closeTmuxPane` mock.module test (reproduces on dev without this
change; isolated run passes).
- `bun run typecheck` -- clean.
- `bun run build` -- clean (esm bundle + tsc + schema).
- Manual harness `.debugging/manual-qa.ts` (uncommitted) drives the same
shape as the real hook chain and resolves the hang scenario in 51 ms.
- Replace the simple scheduledDelays array with an activeTimers Map so
that clearTimeout removes timers from the tracked set.
- This prevents false positives when internal withDispatchTimeout calls
setTimeout for safety timeouts that are immediately cancelled.
- Keeps the test intent unchanged: only genuinely scheduled retries are
counted as delayed duplicate retries.
- Wrap isSessionActive in withDispatchTimeout (capped at 5s) so a
stuck OpenCode SDK status() call cannot block internal prompts forever.
- Catch the timeout and treat session as inactive so the prompt can
proceed rather than hanging indefinitely.
- Add regression test: session.status that never resolves now times out
and allows dispatch instead of hanging the test (and production).
Refs: AGENTS.md internal-message-injection safety note
Red: a43215f24 introduced plugin/build-team-idle-wake-hint-client.ts which
accesses session.promptAsync for method binding. The audit test flagged it
as a raw prompt route offender, breaking CI on dev.
Green: Add the narrow client facade to RAW_PROMPT_ALLOWLIST with the same
justification pattern used for event.ts and recover-unavailable-tool.ts.
The facade binds SDK methods back to the Session instance and performs
no direct dispatch itself; all downstream calls flow through the shared
prompt-async gate.
Verification: bun test src/shared/prompt-async-route-audit.test.ts passes
(6 pass, 0 fail, offenders list empty).
The team-mode wiring at `createEventHandler` extracted
`pluginContext.client.session.promptAsync` and `.status` into a fresh
wrapper object. The methods were copied by reference, so the prompt-async
gate's `session.promptAsync.bind(session)` was binding to that plain
wrapper rather than the underlying SDK `Session` instance. The opencode
SDK's `promptAsync` reads `this._client.post(...)`, so production calls
threw `TypeError: undefined is not an object (evaluating 'this._client')`
on the very first dispatch — fingerprinted in /tmp/oh-my-opencode.log as
688 `background-agent-parent-wake`, 47 `model-suggestion-retry`, and 4
`team-idle-wake-hint` failures over the past three days.
Move the wrapper construction into `buildTeamIdleWakeHintClient`, which
preserves the narrow factory contract while binding both methods back to
the SDK `Session` so `_client` survives the dispatch. Cover the contract
with four BDD-style tests including the historical destructure-only
failure mode so any future regression is caught at unit-test time.
The rule scan cache stored a path[] keyed by (projectRoot|startDir|
skipClaudeUserRules), so two issues stacked up on every tracked tool
call:
- Cache hits still ran safeRealpathSync(realpathSync) and re-derived
isGlobal / distance / isSingleFile for every cached path. That is a
per-candidate sync syscall plus repeated string-prefix walks.
- Sibling files in the same project landed under different startDir
keys, so the entire walk-and-recursive-scan chain repeated even
though every ancestor rule directory was identical.
Store the full RuleFileCandidate[] in the per-call cache so a cache hit
returns immediately with no realpath syscall. Add a separate per-
directory scan cache (getDirScan/setDirScan) keyed by absolute rule
directory path, so two sibling files reuse the same readdir + realpath
work for every shared ancestor.
Microbench (200 files / 20 modules / cached session):
- single sweep: 41.8ms -> 2.5ms (16x)
- 3-pass replay: 88.6ms -> 3.2ms (28x)
Pin the new invariants with two new tests:
- 'does not re-resolve symlinked rule path on cache hit' via a
retargeted directory symlink.
- 'reuses ancestor directory scan for sibling files in the same
project' by deleting the source rule file between the two calls.