Skip to content

ci(windows): dev push CI fails a rotating Windows test by timeout on every run, blocking the publish gate #8324

Description

@code-yeongyu

Summary

Every push-CI run on dev today fails one to three Windows jobs, each time on a different test that hits its timeout (20 s, 30 s or 60 s budgets) while the same test passes on the next run. The rotating set means no single fix lands the gate: publish.yml's gate-reuse needs a fully green ci.yml run on the release-state SHA, so a release currently depends on gh run rerun --failed luck. This is the same shape as #8250 (closed 2026-09-14) and #8294 (closed 2026-09-15) with a fresh rotation.

Evidence (dev push CI, 2026-09-15)

run head job failing test budget
34980106383 820b02f test (windows-latest, 2/2) member injection residency > parked then messaged by task id > detach_rpc resumes one acknowledged turn 20 s (#8316)
34982160042 633a194 test (windows-latest, 1/2) one memory identity across real Bun processes > two synchronized tool writers > twenty linear commits land with observed lock contention 30 s
34982160042 633a194 test (windows-latest, 2/2) #given a console-less parent #when the default RPC child starts #then windowsHide suppresses its console 60 s
34984871235 c5b7e6a senpi-compatibility (windows-latest) ordered delivery mailbox > persists pending messages and enforces count and byte caps 15 s
34984871235 c5b7e6a test (windows-latest, 2/2) omob refresh integration > Windows target and command omob-custom.EXE > executable suffix retained 60 s
34984871235 c5b7e6a test (windows-latest, 2/2) the actual metadata step validates before npm reads and forwards channel outputs 11 s

Each job otherwise passes thousands of tests (e.g. 9318 pass / 1 fail, 6003 pass / 2 fail, 3466 pass / 1 fail). PR-CI Windows legs run a reduced set in about 20 s and stay green, so the rotation is only visible on dev push runs.

The one deterministic Windows failure in the same window, the ulw-loop footer goal cache (#8321), is fixed by #8322 and passes at c5b7e6a.

Expected

  • Tests that spawn real processes or contend on file locks on Windows either get a budget derived from a measured Windows baseline or a deterministic wait on the event they need, so a loaded runner does not turn them into timeouts.
  • A Windows push-CI run on dev is green without reruns for a week of normal merge traffic, measured from the run list.
  • Any test that cannot be made deterministic on the hosted Windows runner is moved to a Windows-specific lane with retries declared in the workflow (bun test --retry), never skipped.

Scope

In: the six tests above and the residency case owned by #8316 / #8318, plus the workflow-level retry policy for the Windows lanes.
Out: PR-CI (already green), non-Windows legs.

Related

#8250, #8294, #8316, #8318, #8321, #8322, #7717 (release gate vs cancelled CI).

Activity

  1. code-yeongyu commented on Sep 16, 2026

    @code-yeongyu
    OwnerAuthor

    Another occurrence, this time on the release-state PR for 5.0.0-beta.64 (#8353), blocking the publish gate exactly as described here:

    • Run 35051767333, job test (windows-latest, 2/2): (fail) member injection residency > #given configured TTL and a real process member #when parked then messaged by task id #then detach_rpc resumes one acknowledged turn [20001.95ms] — 110 pass / 1 fail / 1 skip across 10 files.
    • The adoption PR chore(deps): adopt senpi 2026.9.16 and draft the beta.64 release notes #8352 (same source tree minus the version stamps, same senpi 2026.9.16 pin) passed the same shard minutes earlier, and the concurrent dev push run at 03:32Z failed the same shard, so this is the rotating timeout class, not the release change.
    • Unblocking with gh run rerun --failed on the release-state run; the root cause stays with this issue.
  2. code-yeongyu commented on Sep 16, 2026

    @code-yeongyu
    OwnerAuthor

    Data point from the #8323 remediation (PR #8391, merged ee17c22):

    With the six #8323 members fixed or quarantined-with-evidence, the remaining Windows timeout that surfaced in the PR's soak is member injection residency > … detach_rpc resumes one acknowledged turn — 20003.94 ms on one soak iteration (run 35110047990, full-shard-2), green in all three consecutive PR runs (35113091722, 35116025227, 35117955798). It became visible only because #8391 restored residency.test.ts to the soak's full-shard-2 list, where it had previously run in neither phase. It is the #8213/#8217 marginal residency test already carrying its own quarantine evidence; it is the next Windows budget to attack under this issue.

  3. code-yeongyu commented on Sep 16, 2026

    @code-yeongyu
    OwnerAuthor

    Another instance on the beta.68 release-state PR (#8399), CI run 35136328378, job test (windows-latest, 2/2):

    (fail) install-codex MCP manifest > #given codex installer #when installing omo #then caches research MCPs without ast-grep MCP [30002.67ms]
      ^ this test timed out after 30000ms.
     1 fail
    Ran 128 tests across 12 files. [320.13s]
    

    Same shape as the rotating timeouts tracked here: one test in a Windows shard hits its 30 s budget while the other 127 pass, and prepare-release-state refused the publish. Rerunning the failed job to unblock the release; the durable fix stays with this issue.

  4. code-yeongyu commented on Sep 18, 2026

    @code-yeongyu
    OwnerAuthor

    Another instance, dev push CI run 35290232220 (head 47ad4da1, 2026-09-18), job test (windows-latest, 2/2):

    error: Residency test exceeded its 20000ms budget; cold-revive stages=[{"stage":"setup","elapsed_ms":1.703,...
    (fail) member injection residency > #given configured TTL and a real process member #when parked ...
    error: Cold revival failed; cold-revive stages=[...]
    

    Same 20 s residency budget as the 2026-09-15 row. The shard aborted on it before reaching the two #8427 marker tests, so on that run the residency timeout was the only red. The next dev run on the same tree (release-state #8441's CI) passed residency and failed only #8427's pair — consistent with the rotation this issue describes rather than a deterministic break.

  5. code-yeongyu commented on Sep 18, 2026

    @code-yeongyu
    OwnerAuthor

    member injection residency — already under serial quarantine; recording the measurement

    Following up my 2026-09-18 instance report (run 35290232220, head 47ad4da1). Checking what the pipeline actually does with this test before proposing anything:

    • bunfig.win2.parallel.toml lists packages/senpi-task/src/team/member-extension/residency.test.ts under pathIgnorePatterns, so it is excluded from the --parallel leg.
    • .github/workflows/ci.yml's shard-2-quarantine invocation names it explicitly, so it runs serially, with the same --timeout 20000 budget.

    So the repo's serial-quarantine rule is already applied to it; it is not running under parallel contention. It still timed out on 47ad4da1, which means the remaining failures are inside the serial leg, where the 20 s budget covers a real cold-revive against a spawned child process (realColdRevive("process", …) with idleTimeoutMs: 37).

    What that rules out: "it is starved by the parallel shard" is not the explanation for the remaining instances, because it does not run there.

    What I did not establish: whether 20 s is genuinely too tight for a serial cold-revive on a loaded windows-latest runner, or whether there is a real stall in the revive path. I could not measure it locally — the residency fixture needs a full workspace install and the diag worktree lacks one (Cannot find module '@oh-my-opencode/rules-engine'), and the cost that matters is the runner's, not this machine's.

    The honest next step for whoever picks this up: instrument createColdReviveTrace()'s stage timings on the Windows runner (the trace already carries per-stage elapsed_ms; the failure output prints them) and compare the serial-leg stage profile against the budget. The failing output on 47ad4da1 already shows stages=[{"stage":"setup","elapsed_ms":1.703,...}] truncated — capturing the full array on a red run is the measurement that decides budget-vs-stall.

    Leaving this issue open as the tracker for that; no bare rerun was used to make any run green.

  6. code-yeongyu commented on Sep 18, 2026

    @code-yeongyu
    OwnerAuthor

    member injection residency: measured — it is neither starvation nor a tight budget

    I said the deciding measurement was the full createColdReviveTrace() stage array on a red Windows run. Here it is, from run 35311673965, job test (windows-latest, 2/2), COLD_REVIVE_STAGE lines in the raw log.

    t (ms) delta stage
    1.8 +1.8 setup
    137.4 +135.5 connect
    143.2 +5.8 process_spawned
    2395.3 +2228.8 child bootstrap
    2415.0 +19.7 child provider_registered
    3643.4 +1228.4 child session_start
    4439.6 +703.2 switch_session_ack
    4444.3 +4.4 park_requested
    4547.8 +38.8 process_closed
    4555.9 +3.5 task_send_requested
    4567.0 +5.5 connect (revive)
    4571.0 +4.0 process_spawned (revive)
    5588.8 +993.3 child bootstrap (revive)
    6275.2 +661.5 child session_start (revive)

    Then nothing. Wall clock: the last stage is logged at 05:46:52.9275289Z, and the next line in the log is killed 1 dangling process at 05:47:06.6529978Z — 13.7 seconds of complete silence, ending only when the runner tore the child down at the 20 s deadline.

    So:

    • Not a tight budget. Every piece of real work finishes in 6.3 s. The budget is roughly three times the measured cost, and raising it would only lengthen the silence.
    • Not parallel starvation. The test does not run in the parallel leg at all: it is in pathIgnorePatterns in bunfig.win2.parallel.toml and named in shard-2-quarantine, so it runs serially. The bootstrap segments also show real CPU (user 1.56 s / system 1.41 s), not an I/O-blocked thread.
    • What it actually is: the revived member child starts its session and the turn the test awaits never arrives. A lost or never-delivered message after session_start, on Windows only.

    That reclassifies this from a flaky-timeout to a hang, so a rerun can never be the answer and neither can a bigger budget.

    Likely the same root cause as the win32 socket contract work. #8446 repairs the fake host so it speaks the engine's win32 named-pipe + <path>.secret contract; the awaited turn on the revived session travels over exactly that transport. I have not proven the link, so I am not claiming it fixed — I am recording that this measurement should be re-taken once #8446 lands, before anyone spends more on it.

    Evidence: run 35311673965 (PR #8446's full-matrix Windows shard).

  7. code-yeongyu commented on Sep 20, 2026

    @code-yeongyu
    OwnerAuthor

    Recurred on today's release chain and blocked the publish gate again, exactly as this issue describes.

    Occurrence

    • Run 35525423457, dev @ ec61a9136 (the senpi 2026.9.20 adoption merge)
    • Job: test (windows-latest, 2/2), the only non-green job in the run
    (fail) member injection residency > #given configured TTL and a real process member
    #when parked then messaged by task id #then detach_rpc resumes one acknowledged turn [20007.00ms]
     1 fail
    

    20007.00ms against the --timeout 20000 budget: it ran out of time, it did not assert wrong.

    It is a timeout under load, not a behavioural break

    The child's own staging telemetry in the same log shows where the budget went:

    child_stage bootstrap      elapsed_ms 13653
    child_stage provider_registered elapsed_ms 13668
    child_stage session_start  elapsed_ms 16714
    

    The member process needed 16.7 s just to reach session_start, leaving under 3.3 s for the park / message / detach_rpc round trip the assertion covers. Nothing failed; the runner was too slow to finish in budget.

    Same content passes elsewhere

    The identical tree passed this job in the PR's own CI minutes earlier — run 35525122220, test (windows-latest, 2/2) pass in 16s. Same commit content, 16 s on the PR run versus a 20 s timeout on the dev push. It also passed on the two preceding dev runs (8ad4f4157, 3c0b58c9b).

    Why it matters here

    publish.yml hard-gates on Require successful CI for prepared release source, so this single flaked job blocked the omo release until the job was re-run. That is the cost this issue is tracking: not a broken test, a release that cannot dispatch.

    Related: #8316 closed the revive-hang root cause for this same residency detach_rpc test; what remains here is the budget, since a cold Windows runner can spend the whole 20 s on child bootstrap alone.

  8. code-yeongyu commented on Sep 20, 2026

    @code-yeongyu
    OwnerAuthor

    Second recurrence today, and this time it aborted the omo release.

    Occurrence

    ##[error]Release-state PR #8546 has failing required checks; refusing to publish npm packages.
    
    • Failing job: test (windows-latest, 1/2) (run 35527280944), 15m36s, 4 fail:
    (fail) one memory identity across real Bun processes > #given two synchronized tool writers
           #when each makes ten writes #then twenty linear commits land with observed lock contention [30708.04ms]
    (fail) one memory identity across real Bun processes > #given a writer SIGKILLed while holding
           the shared lock #when its peer continues #then liveness takeover completes twenty clean commits [30255.38ms]
    (fail) recall-wake lock domain: counting FIFO lease > ... never more than two leases live [74.98ms]
    (fail) pre-commit hook > #given valid memory markdown #when it is committed
           #then the hook accepts every supported shape [3466.84ms]
    

    Two of the four are clean 30 s timeouts (30708ms, 30255ms against a 30 s budget). All four are git/fs-heavy memory-identity and lock-domain tests — the same family as #7709.

    Why it is not the release content

    The release-state branch is dev plus version stamping. test (windows-latest, 1/2) passed on dev @ ec61a9136 in run 35525423457 attempt 2, and passed on PR #8545 before that. Version stamping does not touch memory locks or the pre-commit hook.

    Why this issue matters more than the earlier occurrence

    The first one today (recorded above) only cost a job re-run. This one consumed the whole release: prepare-release-state had already stamped 5.0.0-beta.80 across the monorepo and opened the PR, then refused to publish and exited non-zero, leaving the release half-done and needing manual recovery. Two different Windows shards (2/2 then 1/2) on two different rotating tests within one hour, exactly as the title describes.

    The publish gate is correct to refuse on red. The defect is that the red is not real.

  9. code-yeongyu commented on Sep 21, 2026

    @code-yeongyu
    OwnerAuthor

    Recurrence, 2026-09-22: blocking the isolation-core fix that the beta.82 release waits on

    Run 35635164917, job test (windows-latest, 2/2) on PR #8607, attempt 1.

    Every other job on that PR was green (ubuntu 1/2 and 2/2, macos 1/2 and 2/2, windows 1/2, senpi-compatibility on ubuntu and macos). The four annotations on the failing shard are all this family, with no isolation-core failure left:

    AssertionError: Expected values to be strictly equal:
      + 'delivery_uncertain'  - 'revived'
      at packages/senpi-task/src/lifecycle/__fixtures__/real-cold-revive.ts:156:12
    
    error: Cold revival failed
      at packages/senpi-task/src/lifecycle/__fixtures__/real-cold-revive.ts:196:47
    
    error: Residency test exceeded its 20000ms budget
      at packages/senpi-task/src/team/member-extension/residency.test.ts:28:30
    

    Timing from the stage telemetry: the child burned ~3.9 s inside bootstrap alone before provider_registered, against a 20 s test budget.

    Cost this time: the shard is the last red check on #8607, which is itself the fix the 5.0.0-beta.82 release is blocked behind. Attempt 2 is running now.

  10. code-yeongyu commented on Sep 21, 2026

    @code-yeongyu
    OwnerAuthor

    Residual windows-latest 2/2 failures — attribution vs the isolation-core work (PR #8607)

    PR #8607 (merged as 4e894a6) fixed every test #8604 named. On its final head d0613e1, the windows-latest 2/2 shard still fails exactly four tests, all in this issue's family:

    • host-session handle reattach > #given a turn in flight #when the host dies and comes back #then the handle reopens the path and re-prompts a continuation
    • host-session handle reattach > #given an idle child #when the host dies and comes back #then the handle reopens the path without re-prompting
    • RpcHostRunner transport recovery > #given a child mid-turn #when the daemon dies and a new one answers the socket #then the child reopens its path there and is re-prompted once
    • two parents on one daemon > #given a daemon that died under 32 live children #when they park and it comes back #then the next session start reconciles every child

    Dev-run attribution (same family failing before this branch):

    • dev ac6addd — run 35622630569
    • dev 40872be — run 35628124210 (windows 2/2 cancelled in that run; family failed in the earlier full runs)
    • dev 73c6bbc — run 35629230915 job 106431480010
    • this PR's d0613e1 — run 35649636084 job 106498719257

    Full test-by-test comparison: the "Attribution matrix" comment on PR #8607.

    Honest nuance (pAX): on the release base 73c6bbc the pre-existing set is handle-reattach x2 + transport recovery x1 only — two parents on one daemon > ... 32 live children PASSED there and FAILED on ac6addd (pre-#8605) and on this PR's revert head. That one case is therefore intermittent on dev and its "Received 64" on d0613e1 is not fully explained by the base; it is outside #8604's scope (its failure mode — prompts doubled during restart reconcile — is tracked separately in the fake-host handoff issue).

  11. code-yeongyu commented on Sep 21, 2026

    @code-yeongyu
    OwnerAuthor

    The Received 64 is now on dev itself, and it was not there before the #8607 merge

    Before / after on the same job name, test (windows-latest, 2/2):

    Before — dev 73c6bbcb7, run 35629230915, job 106431480010. From the job log:

    (pass) two parents on one daemon > #given a daemon that died under 32 live children
           #when they park and it comes back #then the next session start reconciles every child
    

    After — dev 4e894a6ef (the #8607 merge), run 35651695475, job 106505510155. From the check-run annotations:

    expect(received).toHaveLength(expected)
      Expected length: 32   Received length: 64
      at packages/senpi-task/src/runners/rpc-host.integration.test.ts:161:29
    

    Evidence limit worth stating: the "after" side is annotation-based because that run is still in flight and its job log was not retrievable yet, so I am comparing a log line against an annotation for the same file and line. Line 161 is the case above.

    That makes this particular assertion a regression dev acquired at the #8607 merge, not a member of the pre-existing family - which matters because it was classified the other way while triaging the PR. The other failures in the same job do look like the established family plus one new class:

    listen EADDRINUSE  \\.\pipe\senpi-rpc-<id>                                  x3
    EBUSY: resource busy or locked, rm '...\Temp\omo-worker-compile-<id>'
      at script/senpi-worker-compile.test.ts:51:5
    

    The EBUSY on the worker-compile temp directory is one I have not seen in the earlier runs of this shard.

    Release impact, since this issue is where the shard's release cost is tracked: 5.0.0-beta.82 remains unpublishable. publish.yml gate-reuse needs a ci.yml run at the prepared SHA with conclusion success, and dev does not currently produce one.

  12. code-yeongyu commented on Sep 21, 2026

    @code-yeongyu
    OwnerAuthor

    The new EBUSY looks like a Windows cleanup race, and it is first-seen on this run

    Following up on the EBUSY I listed above, since it is the one failure class in that job I had not seen before.

    error: EBUSY: resource busy or locked, rm '...\Temp\omo-worker-compile-<id>'
      at script/senpi-worker-compile.test.ts:51:5
    

    Line 51 is the finally cleanup:

    const moved = join(relocated, process.platform === "win32" ? "omo.exe" : "omo")
    renameSync(binary, moved)
    rmSync(buildRoot, { recursive: true })
    const result = spawnSync(moved, [], { cwd: relocated, encoding: "utf8", timeout: 10_000 })
    ...
    } finally {
      rmSync(scratch, { recursive: true, force: true })   // <- line 51
    }

    The test executes omo.exe from inside scratch and then removes scratch in the finally. Windows holds the image lock on an executable slightly past process exit, so a removal issued immediately after spawnSync returns can still see the file busy. force: true suppresses ENOENT, not EBUSY.

    Status rather than diagnosis: this is a single occurrence. It is absent from the annotations I captured for the dev run at 73c6bbcb7 and from all four #8607 branch runs, and present once on dev at 4e894a6ef. So I would treat it as sporadic in the same "Windows releases a resource later than the test expects" family as the EADDRINUSE pipe collisions, rather than as a deterministic break, until it shows up a second time.

    Recording it here because it lands on the same shard that gates the release, so it can turn a green-looking fix run red on its own.

  13. code-yeongyu commented on Sep 21, 2026

    @code-yeongyu
    OwnerAuthor

    Cross-reference so this issue is not held responsible for the release block: on the dev run at 4e894a6ef (run 35651695475, job 106505510155) the three rpc-host failures on test (windows-latest, 2/2) — host-session handle reattach x2 and RpcHostRunner transport recovery — each fail with listen EADDRINUSE on the senpi-rpc pipe, not with a 20s timeout. Details and the reasoning are on #8609. Nothing in that job matches this issue's timeout signature.

  14. code-yeongyu commented on Sep 21, 2026

    @code-yeongyu
    OwnerAuthor

    Correction from me: the Expected 32 / Received 64 is not a regression, and it is not this issue's timeout family either

    Two things I wrote on this issue were wrong, and both were wrong in the same way - I read a symptom and named a cause without checking the code delta.

    1. I called the 64 a regression dev acquired at the fix(isolation-core): green the windows-latest suite #8607 merge, comparing 73c6bbcb7 (pass) against 4e894a6ef (fail). The comparison between those two commits touches no file under senpi-task/runners/rpc-host - the fake-host rebind work in that PR was fully undone by its revert, and what landed was isolation-core plus its fakes and spawn test. Unchanged code cannot regress; that was intermittency, which is the same conclusion reached independently from ac6addd6d versus 73c6bbcb7.
    2. My "after" evidence was an annotation snapshot taken while the run was still in flight. The completed log for that run carries no Received: 64 at all.

    The real explanation is now on #8609 with a fix in #8614: the 32-children test crashes the host right after startChildren, which returns once children have started rather than once their opening turns have ended, so platform timing picked which branch of the #8577 S4 revival contract ran. Mid-turn revival re-prompts once per child by contract, which is where 64 comes from.

    Separately, and relevant to this issue's own hypothesis: the three EADDRINUSE failures that sat on this shard were not runner performance either. They were a same-name rebind in the fake host's restart(), fixed in #8613, and their removal took the shard from five failures to one. Worth weighing before attributing further failures here to slow runners.

  15. code-yeongyu commented on Sep 24, 2026

    @code-yeongyu
    OwnerAuthor

    Recurrence on dev push f83066d80 (run 35963717011, test (windows-latest, 2/2)): member injection residency > … detach_rpc resumes one acknowledged turn timed out at 20004.22ms, the only failure in the shard (1 fail). The commit is a senpi pin bump (2026.9.23-5 -> 2026.9.24) that does not touch senpi-task residency. The same file set passed on its PR run (#8782, all checks green). The run was then cancelled by the next dev push.

  16. code-yeongyu commented on Sep 28, 2026

    @code-yeongyu
    OwnerAuthor

    Closing with evidence: dev's Windows CI is green again, and it stays green across reruns.

    C4: dev push CI on a head containing every win-ci fix (e89647d), run 36360065018:

    Fixes that landed (merge commits):

    Related fixes owned by other sessions: #9005/#9021 (child start), #9031 (#9029 cold child import), #9034 (#9030 team lock delete-pending), #9038 (#9036).

    Every fix was soaked 3 times on windows-latest before merge (C1), and the combined set was full-matrix green 3 times before landing (C2). No test was skipped or deleted and no timeout was widened: wall-clock assertions became event or ordering proofs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions