Repository navigation
ci(windows): dev push CI fails a rotating Windows test by timeout on every run, blocking the publish gate #8324
Description
Activity
Another occurrence, this time on the release-state PR for 5.0.0-beta.64 (#8353), blocking the publish gate exactly as described here:
- Run 35051767333, job
test (windows-latest, 2/2):(fail) member injection residency > #given configured TTL and a real process member #when parked then messaged by task id #then detach_rpc resumes one acknowledged turn [20001.95ms]— 110 pass / 1 fail / 1 skip across 10 files. - The adoption PR chore(deps): adopt senpi 2026.9.16 and draft the beta.64 release notes #8352 (same source tree minus the version stamps, same senpi 2026.9.16 pin) passed the same shard minutes earlier, and the concurrent dev push run at 03:32Z failed the same shard, so this is the rotating timeout class, not the release change.
- Unblocking with
gh run rerun --failedon the release-state run; the root cause stays with this issue.
- Run 35051767333, job
Data point from the #8323 remediation (PR #8391, merged ee17c22):
With the six #8323 members fixed or quarantined-with-evidence, the remaining Windows timeout that surfaced in the PR's soak is
member injection residency > … detach_rpc resumes one acknowledged turn— 20003.94 ms on one soak iteration (run 35110047990,full-shard-2), green in all three consecutive PR runs (35113091722, 35116025227, 35117955798). It became visible only because #8391 restoredresidency.test.tsto the soak'sfull-shard-2list, where it had previously run in neither phase. It is the #8213/#8217 marginal residency test already carrying its own quarantine evidence; it is the next Windows budget to attack under this issue.Another instance on the beta.68 release-state PR (#8399), CI run 35136328378, job
test (windows-latest, 2/2):(fail) install-codex MCP manifest > #given codex installer #when installing omo #then caches research MCPs without ast-grep MCP [30002.67ms] ^ this test timed out after 30000ms. 1 fail Ran 128 tests across 12 files. [320.13s]Same shape as the rotating timeouts tracked here: one test in a Windows shard hits its 30 s budget while the other 127 pass, and
prepare-release-staterefused the publish. Rerunning the failed job to unblock the release; the durable fix stays with this issue.Another instance, dev push CI run 35290232220 (head
47ad4da1, 2026-09-18), jobtest (windows-latest, 2/2):error: Residency test exceeded its 20000ms budget; cold-revive stages=[{"stage":"setup","elapsed_ms":1.703,... (fail) member injection residency > #given configured TTL and a real process member #when parked ... error: Cold revival failed; cold-revive stages=[...]Same 20 s residency budget as the 2026-09-15 row. The shard aborted on it before reaching the two #8427 marker tests, so on that run the residency timeout was the only red. The next dev run on the same tree (release-state #8441's CI) passed residency and failed only #8427's pair — consistent with the rotation this issue describes rather than a deterministic break.
member injection residency— already under serial quarantine; recording the measurementFollowing up my 2026-09-18 instance report (run 35290232220, head
47ad4da1). Checking what the pipeline actually does with this test before proposing anything:bunfig.win2.parallel.tomllistspackages/senpi-task/src/team/member-extension/residency.test.tsunderpathIgnorePatterns, so it is excluded from the--parallelleg..github/workflows/ci.yml'sshard-2-quarantineinvocation names it explicitly, so it runs serially, with the same--timeout 20000budget.
So the repo's serial-quarantine rule is already applied to it; it is not running under parallel contention. It still timed out on
47ad4da1, which means the remaining failures are inside the serial leg, where the 20 s budget covers a real cold-revive against a spawned child process (realColdRevive("process", …)withidleTimeoutMs: 37).What that rules out: "it is starved by the parallel shard" is not the explanation for the remaining instances, because it does not run there.
What I did not establish: whether 20 s is genuinely too tight for a serial cold-revive on a loaded
windows-latestrunner, or whether there is a real stall in the revive path. I could not measure it locally — the residency fixture needs a full workspace install and the diag worktree lacks one (Cannot find module '@oh-my-opencode/rules-engine'), and the cost that matters is the runner's, not this machine's.The honest next step for whoever picks this up: instrument
createColdReviveTrace()'s stage timings on the Windows runner (the trace already carries per-stageelapsed_ms; the failure output prints them) and compare the serial-leg stage profile against the budget. The failing output on47ad4da1already showsstages=[{"stage":"setup","elapsed_ms":1.703,...}]truncated — capturing the full array on a red run is the measurement that decides budget-vs-stall.Leaving this issue open as the tracker for that; no bare rerun was used to make any run green.
member injection residency: measured — it is neither starvation nor a tight budgetI said the deciding measurement was the full
createColdReviveTrace()stage array on a red Windows run. Here it is, from run 35311673965, jobtest (windows-latest, 2/2),COLD_REVIVE_STAGElines in the raw log.t (ms) delta stage 1.8 +1.8 setup 137.4 +135.5 connect 143.2 +5.8 process_spawned 2395.3 +2228.8 child bootstrap 2415.0 +19.7 child provider_registered 3643.4 +1228.4 child session_start 4439.6 +703.2 switch_session_ack 4444.3 +4.4 park_requested 4547.8 +38.8 process_closed 4555.9 +3.5 task_send_requested 4567.0 +5.5 connect (revive) 4571.0 +4.0 process_spawned (revive) 5588.8 +993.3 child bootstrap (revive) 6275.2 +661.5 child session_start (revive) Then nothing. Wall clock: the last stage is logged at
05:46:52.9275289Z, and the next line in the log iskilled 1 dangling processat05:47:06.6529978Z— 13.7 seconds of complete silence, ending only when the runner tore the child down at the 20 s deadline.So:
- Not a tight budget. Every piece of real work finishes in 6.3 s. The budget is roughly three times the measured cost, and raising it would only lengthen the silence.
- Not parallel starvation. The test does not run in the parallel leg at all: it is in
pathIgnorePatternsinbunfig.win2.parallel.tomland named inshard-2-quarantine, so it runs serially. The bootstrap segments also show real CPU (user1.56 s /system1.41 s), not an I/O-blocked thread. - What it actually is: the revived member child starts its session and the turn the test awaits never arrives. A lost or never-delivered message after
session_start, on Windows only.
That reclassifies this from a flaky-timeout to a hang, so a rerun can never be the answer and neither can a bigger budget.
Likely the same root cause as the win32 socket contract work. #8446 repairs the fake host so it speaks the engine's win32 named-pipe +
<path>.secretcontract; the awaited turn on the revived session travels over exactly that transport. I have not proven the link, so I am not claiming it fixed — I am recording that this measurement should be re-taken once #8446 lands, before anyone spends more on it.Evidence: run 35311673965 (PR #8446's full-matrix Windows shard).
Recurred on today's release chain and blocked the publish gate again, exactly as this issue describes.
Occurrence
- Run 35525423457,
dev@ec61a9136(the senpi 2026.9.20 adoption merge) - Job:
test (windows-latest, 2/2), the only non-green job in the run
(fail) member injection residency > #given configured TTL and a real process member #when parked then messaged by task id #then detach_rpc resumes one acknowledged turn [20007.00ms] 1 fail20007.00msagainst the--timeout 20000budget: it ran out of time, it did not assert wrong.It is a timeout under load, not a behavioural break
The child's own staging telemetry in the same log shows where the budget went:
child_stage bootstrap elapsed_ms 13653 child_stage provider_registered elapsed_ms 13668 child_stage session_start elapsed_ms 16714The member process needed 16.7 s just to reach
session_start, leaving under 3.3 s for the park / message / detach_rpc round trip the assertion covers. Nothing failed; the runner was too slow to finish in budget.Same content passes elsewhere
The identical tree passed this job in the PR's own CI minutes earlier — run 35525122220,
test (windows-latest, 2/2)pass in 16s. Same commit content, 16 s on the PR run versus a 20 s timeout on the dev push. It also passed on the two preceding dev runs (8ad4f4157,3c0b58c9b).Why it matters here
publish.ymlhard-gates onRequire successful CI for prepared release source, so this single flaked job blocked the omo release until the job was re-run. That is the cost this issue is tracking: not a broken test, a release that cannot dispatch.Related: #8316 closed the revive-hang root cause for this same
residency detach_rpctest; what remains here is the budget, since a cold Windows runner can spend the whole 20 s on child bootstrap alone.- Run 35525423457,
Second recurrence today, and this time it aborted the omo release.
Occurrence
- Release run 35527200065 (
release 5.0.0-beta.80) - It opened release-state PR release: v5.0.0-beta.80 #8546, polled 32 times over ~17 minutes, then:
##[error]Release-state PR #8546 has failing required checks; refusing to publish npm packages.- Failing job:
test (windows-latest, 1/2)(run 35527280944), 15m36s, 4 fail:
(fail) one memory identity across real Bun processes > #given two synchronized tool writers #when each makes ten writes #then twenty linear commits land with observed lock contention [30708.04ms] (fail) one memory identity across real Bun processes > #given a writer SIGKILLed while holding the shared lock #when its peer continues #then liveness takeover completes twenty clean commits [30255.38ms] (fail) recall-wake lock domain: counting FIFO lease > ... never more than two leases live [74.98ms] (fail) pre-commit hook > #given valid memory markdown #when it is committed #then the hook accepts every supported shape [3466.84ms]Two of the four are clean 30 s timeouts (
30708ms,30255msagainst a 30 s budget). All four are git/fs-heavy memory-identity and lock-domain tests — the same family as #7709.Why it is not the release content
The release-state branch is
devplus version stamping.test (windows-latest, 1/2)passed ondev@ec61a9136in run 35525423457 attempt 2, and passed on PR #8545 before that. Version stamping does not touch memory locks or the pre-commit hook.Why this issue matters more than the earlier occurrence
The first one today (recorded above) only cost a job re-run. This one consumed the whole release:
prepare-release-statehad already stamped5.0.0-beta.80across the monorepo and opened the PR, then refused to publish and exited non-zero, leaving the release half-done and needing manual recovery. Two different Windows shards (2/2then1/2) on two different rotating tests within one hour, exactly as the title describes.The publish gate is correct to refuse on red. The defect is that the red is not real.
- Release run 35527200065 (
Recurrence, 2026-09-22: blocking the isolation-core fix that the beta.82 release waits on
Run 35635164917, job
test (windows-latest, 2/2)on PR #8607, attempt 1.Every other job on that PR was green (ubuntu 1/2 and 2/2, macos 1/2 and 2/2, windows 1/2, senpi-compatibility on ubuntu and macos). The four annotations on the failing shard are all this family, with no isolation-core failure left:
AssertionError: Expected values to be strictly equal: + 'delivery_uncertain' - 'revived' at packages/senpi-task/src/lifecycle/__fixtures__/real-cold-revive.ts:156:12 error: Cold revival failed at packages/senpi-task/src/lifecycle/__fixtures__/real-cold-revive.ts:196:47 error: Residency test exceeded its 20000ms budget at packages/senpi-task/src/team/member-extension/residency.test.ts:28:30Timing from the stage telemetry: the child burned ~3.9 s inside
bootstrapalone beforeprovider_registered, against a 20 s test budget.Cost this time: the shard is the last red check on #8607, which is itself the fix the
5.0.0-beta.82release is blocked behind. Attempt 2 is running now.Residual windows-latest 2/2 failures — attribution vs the isolation-core work (PR #8607)
PR #8607 (merged as 4e894a6) fixed every test #8604 named. On its final head d0613e1, the windows-latest 2/2 shard still fails exactly four tests, all in this issue's family:
host-session handle reattach > #given a turn in flight #when the host dies and comes back #then the handle reopens the path and re-prompts a continuationhost-session handle reattach > #given an idle child #when the host dies and comes back #then the handle reopens the path without re-promptingRpcHostRunner transport recovery > #given a child mid-turn #when the daemon dies and a new one answers the socket #then the child reopens its path there and is re-prompted oncetwo parents on one daemon > #given a daemon that died under 32 live children #when they park and it comes back #then the next session start reconciles every child
Dev-run attribution (same family failing before this branch):
- dev ac6addd — run 35622630569
- dev 40872be — run 35628124210 (windows 2/2 cancelled in that run; family failed in the earlier full runs)
- dev 73c6bbc — run 35629230915 job 106431480010
- this PR's d0613e1 — run 35649636084 job 106498719257
Full test-by-test comparison: the "Attribution matrix" comment on PR #8607.
Honest nuance (pAX): on the release base 73c6bbc the pre-existing set is
handle-reattachx2 +transport recoveryx1 only —two parents on one daemon > ... 32 live childrenPASSED there and FAILED on ac6addd (pre-#8605) and on this PR's revert head. That one case is therefore intermittent on dev and its "Received 64" on d0613e1 is not fully explained by the base; it is outside #8604's scope (its failure mode — prompts doubled during restart reconcile — is tracked separately in the fake-host handoff issue).The
Received 64is now ondevitself, and it was not there before the #8607 mergeBefore / after on the same job name,
test (windows-latest, 2/2):Before — dev
73c6bbcb7, run 35629230915, job 106431480010. From the job log:(pass) two parents on one daemon > #given a daemon that died under 32 live children #when they park and it comes back #then the next session start reconciles every childAfter — dev
4e894a6ef(the #8607 merge), run 35651695475, job 106505510155. From the check-run annotations:expect(received).toHaveLength(expected) Expected length: 32 Received length: 64 at packages/senpi-task/src/runners/rpc-host.integration.test.ts:161:29Evidence limit worth stating: the "after" side is annotation-based because that run is still in flight and its job log was not retrievable yet, so I am comparing a log line against an annotation for the same file and line. Line 161 is the case above.
That makes this particular assertion a regression
devacquired at the #8607 merge, not a member of the pre-existing family - which matters because it was classified the other way while triaging the PR. The other failures in the same job do look like the established family plus one new class:listen EADDRINUSE \\.\pipe\senpi-rpc-<id> x3 EBUSY: resource busy or locked, rm '...\Temp\omo-worker-compile-<id>' at script/senpi-worker-compile.test.ts:51:5The
EBUSYon the worker-compile temp directory is one I have not seen in the earlier runs of this shard.Release impact, since this issue is where the shard's release cost is tracked:
5.0.0-beta.82remains unpublishable.publish.ymlgate-reuse needs aci.ymlrun at the prepared SHA with conclusion success, anddevdoes not currently produce one.The new
EBUSYlooks like a Windows cleanup race, and it is first-seen on this runFollowing up on the
EBUSYI listed above, since it is the one failure class in that job I had not seen before.error: EBUSY: resource busy or locked, rm '...\Temp\omo-worker-compile-<id>' at script/senpi-worker-compile.test.ts:51:5Line 51 is the
finallycleanup:const moved = join(relocated, process.platform === "win32" ? "omo.exe" : "omo") renameSync(binary, moved) rmSync(buildRoot, { recursive: true }) const result = spawnSync(moved, [], { cwd: relocated, encoding: "utf8", timeout: 10_000 }) ... } finally { rmSync(scratch, { recursive: true, force: true }) // <- line 51 }
The test executes
omo.exefrom insidescratchand then removesscratchin thefinally. Windows holds the image lock on an executable slightly past process exit, so a removal issued immediately afterspawnSyncreturns can still see the file busy.force: truesuppresses ENOENT, not EBUSY.Status rather than diagnosis: this is a single occurrence. It is absent from the annotations I captured for the
devrun at73c6bbcb7and from all four #8607 branch runs, and present once ondevat4e894a6ef. So I would treat it as sporadic in the same "Windows releases a resource later than the test expects" family as theEADDRINUSEpipe collisions, rather than as a deterministic break, until it shows up a second time.Recording it here because it lands on the same shard that gates the release, so it can turn a green-looking fix run red on its own.
Cross-reference so this issue is not held responsible for the release block: on the
devrun at4e894a6ef(run 35651695475, job 106505510155) the three rpc-host failures ontest (windows-latest, 2/2)—host-session handle reattachx2 andRpcHostRunner transport recovery— each fail withlisten EADDRINUSEon thesenpi-rpcpipe, not with a 20s timeout. Details and the reasoning are on #8609. Nothing in that job matches this issue's timeout signature.Correction from me: the
Expected 32 / Received 64is not a regression, and it is not this issue's timeout family eitherTwo things I wrote on this issue were wrong, and both were wrong in the same way - I read a symptom and named a cause without checking the code delta.
- I called the 64 a regression
devacquired at the fix(isolation-core): green the windows-latest suite #8607 merge, comparing73c6bbcb7(pass) against4e894a6ef(fail). The comparison between those two commits touches no file undersenpi-task/runners/rpc-host- the fake-host rebind work in that PR was fully undone by its revert, and what landed was isolation-core plus its fakes and spawn test. Unchanged code cannot regress; that was intermittency, which is the same conclusion reached independently fromac6addd6dversus73c6bbcb7. - My "after" evidence was an annotation snapshot taken while the run was still in flight. The completed log for that run carries no
Received: 64at all.
The real explanation is now on #8609 with a fix in #8614: the 32-children test crashes the host right after
startChildren, which returns once children have started rather than once their opening turns have ended, so platform timing picked which branch of the #8577 S4 revival contract ran. Mid-turn revival re-prompts once per child by contract, which is where 64 comes from.Separately, and relevant to this issue's own hypothesis: the three
EADDRINUSEfailures that sat on this shard were not runner performance either. They were a same-name rebind in the fake host'srestart(), fixed in #8613, and their removal took the shard from five failures to one. Worth weighing before attributing further failures here to slow runners.- I called the 64 a regression
Recurrence on dev push
f83066d80(run 35963717011,test (windows-latest, 2/2)):member injection residency > … detach_rpc resumes one acknowledged turntimed out at 20004.22ms, the only failure in the shard (1 fail). The commit is a senpi pin bump (2026.9.23-5 -> 2026.9.24) that does not touch senpi-task residency. The same file set passed on its PR run (#8782, all checks green). The run was then cancelled by the next dev push.Closing with evidence: dev's Windows CI is green again, and it stays green across reruns.
C4: dev push CI on a head containing every win-ci fix (e89647d), run 36360065018:
- attempt 1: all Windows jobs green (
test (windows-latest, 1/2),test (windows-latest, 2/2),senpi-compatibility (windows-latest),codex-compatibility (windows-latest, platform)), and the test steps actually executed - attempt 2: all green, same jobs
- attempt 3: all green, same jobs
Also green: dev push 36358896891 attempt 1 on 6f65ada (the fix(windows-ci): land the win-ci fixes (kibitzer ticket sharing, mailbox journal, deny ordering, applyEdit, parent wake, downloader probe, git bound, test removals) #8961 merge).
Fixes that landed (merge commits):
- fix(senpi-task): the liveness probe reaches the Windows daemon over its named pipe #8971 (8c7cf89): [Bug]: DAG nodes stay running after their child task records complete, blocking dependents #8932 foreign-owner liveness over the named pipe, with taskkill as the authority (dev Windows CI red since #8935: #8932 daemon-child liveness tests fail on windows-latest #8942)
- test(omo-senpi): real win32 dead checker for the #6396 stdin-pipe test #9033 (20f9852): real win32 dead checker for the [Bug]: Senpi crashes with unhandled EPIPE when comment-checker exits before consuming stdin #6396 stdin-pipe test after fix(comment-checker): replace a stale cached checker and report one that cannot run #8877 (Windows CI: comment-checker stdin-pipe test (#6396) fakes the checker with bun itself and fails under #8877's exit handling #9032)
- fix(windows-ci): land the win-ci fixes (kibitzer ticket sharing, mailbox journal, deny ordering, applyEdit, parent wake, downloader probe, git bound, test removals) #8961 (6f65ada): kibitzer ticket sharing (Windows ticket-sharing EPERM rejects Kibitzer wake waiters and leaks test work #8953), mailbox journal (fix(omo-senpi): avoid whole-queue mailbox rewrites #8948), deny ordering proof (Windows CI flakes on HostSessionClient deny latency assertion #8950), applyEdit subscribe-first (ci(windows): LSP applyEdit lease test polls past its response budget #8954), parent-wake single hold (Parent wake that meets the promptAsync hold waits a whole extra hold (Windows CI: parent-wake empty-turn recovery test) #8951), downloader single-process probe (Windows CI: comment-checker downloader test cold-starts six probe processes and one hits its 10 s abort (exit 143) #8979), git stall bound (isolation-core: nested baseline can hang indefinitely when Git stalls #8997), and removeTree for Bun's ignored fs.rm retries (Windows CI: test cleanups rely on fs.rm maxRetries, which Bun ignores (EBUSY teardown failures) #9025)
- test(isolation-core): event-ordered proof for the git alias-survivor settle (no wall clock) #9040 (e89647d): event-ordered git alias-survivor proof (Windows CI: isolation-core git survivor test bounds a causal property with a 5s wall clock (6093 ms under load) #9039)
Related fixes owned by other sessions: #9005/#9021 (child start), #9031 (#9029 cold child import), #9034 (#9030 team lock delete-pending), #9038 (#9036).
Every fix was soaked 3 times on windows-latest before merge (C1), and the combined set was full-matrix green 3 times before landing (C2). No test was skipped or deleted and no timeout was widened: wall-clock assertions became event or ordering proofs.
- attempt 1: all Windows jobs green (
Summary
Every push-CI run on
devtoday fails one to three Windows jobs, each time on a different test that hits its timeout (20 s, 30 s or 60 s budgets) while the same test passes on the next run. The rotating set means no single fix lands the gate:publish.yml'sgate-reuseneeds a fully greenci.ymlrun on the release-state SHA, so a release currently depends ongh run rerun --failedluck. This is the same shape as #8250 (closed 2026-09-14) and #8294 (closed 2026-09-15) with a fresh rotation.Evidence (dev push CI, 2026-09-15)
Each job otherwise passes thousands of tests (e.g. 9318 pass / 1 fail, 6003 pass / 2 fail, 3466 pass / 1 fail). PR-CI Windows legs run a reduced set in about 20 s and stay green, so the rotation is only visible on
devpush runs.The one deterministic Windows failure in the same window, the ulw-loop footer goal cache (#8321), is fixed by #8322 and passes at c5b7e6a.
Expected
devis green without reruns for a week of normal merge traffic, measured from the run list.bun test --retry), never skipped.Scope
In: the six tests above and the residency case owned by #8316 / #8318, plus the workflow-level retry policy for the Windows lanes.
Out: PR-CI (already green), non-Windows legs.
Related
#8250, #8294, #8316, #8318, #8321, #8322, #7717 (release gate vs cancelled CI).