Repository navigation
fix(coding-agent): resolve v0.18.8 CI lifecycle regressions - #6512
Yeachan-Heo wants to merge 38 commits into
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 44e5d54785
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| using tempDir = TempDir.createSync("@gjc-python-lifecycle-redteam-"); | ||
| const controller = new AbortController(); | ||
| const listeners = countAbortListeners(controller.signal); | ||
| let listeners: ReturnType<typeof countAbortListeners> | undefined; |
There was a problem hiding this comment.
Replace the forbidden
ReturnType utility
Replace this inferred utility type with a named type for the listener counter result; the repository contract explicitly prohibits every use of ReturnType<>, including in tests.
AGENTS.md reference: AGENTS.md:L127-L127
Useful? React with 👍 / 👎.
probepark
left a comment
There was a problem hiding this comment.
Review (head 44e5d54, gajae-reviewer on behalf of probepark)
CI: PR cause 1 — Affected path validation / test:packages/coding-agent/src/sdk/host/session-runtime.test.ts fails with 2 tests; the exact-head run is still in progress with the shard pending.
Scope: +161 / -31, 9 files — coding-agent async job management, SDK session runtime, lifecycle/lease/ACP regression tests, changelog.
Conventions: CHANGELOG fragment present, generated files none, labels none.
ocr: blocking 0 / nit 0
Notable:
packages/coding-agent/src/sdk/host/session-runtime.test.ts:6178— the exact-head CI run reportsterminalizes a synchronous throw during a todo-reminder continuationfailing because 2agent_endframes are emitted where the test requires 1.packages/coding-agent/src/sdk/host/session-runtime.test.ts:5392— the exact-head CI run also reportsa captured cancelled end stays private and recoverable while tools are unproventiming out without a terminal reconciliation status.
Blocking: The exact-head affected test job fails for two lifecycle cases above. This must be fixed and rerun before approval.
Verdict: gaja.pr-review-verdict.v1 needs-human sha256:227fad5f72b712424d9ea21c15cf1428e26ae9083e12b76940efe75e90fbedd6 reviewer:critic reviewer-id:gajae-reviewer evidence:ci-pr-failure;session-runtime-lifecycle-tests-failing;ocr-checked
PR body verdict line count=0, not updated. Suggested verdict line: gaja.pr-review-verdict.v1 needs-human sha256:227fad5f72b712424d9ea21c15cf1428e26ae9083e12b76940efe75e90fbedd6 reviewer:critic reviewer-id:gajae-reviewer evidence:ci-pr-failure;session-runtime-lifecycle-tests-failing;ocr-checked
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d26f93421a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| rec.queued?.ownerId !== entry.ownerId || | ||
| rec.queued?.seq !== entry.seq | ||
| ) { | ||
| this.#resumeQueue.splice(index, 1); |
There was a problem hiding this comment.
Settle stale queue entries instead of silently dropping them
When a queued resume is replaced by a record for the same subagent ID—the scenario covered by the new replacement-owner test—this mismatch branch removes the old entry without publishing a terminal event for queued:<id>:<seq> or unregistering the owned registration created when resumeToolCallId was supplied. A waiter that captured the old generation can therefore remain unresolved, while owned abort/deadline settlement continues to see the orphaned registration as unsettled. Retire the exact stale entry's registration and terminalize its generation without mutating the replacement record before removing it.
Useful? React with 👍 / 👎.
probepark
left a comment
There was a problem hiding this comment.
Review (head d26f934, gajae-reviewer on behalf of probepark)
CI: PR 원인 1개 — Affected path validation / test:packages/coding-agent/src/sdk/host/session-runtime.test.ts fails because the changed accepted-control zero-execution bound (#4668) > a captured cancelled end stays private and recoverable while tools are unproven test never reaches a terminal reconciliation status (turn.prompt_status never reported a terminal reconciliation status, run https://github.com/Yeachan-Heo/gajae-code/actions/runs/37773298855/job/113303899006).
범위: +169 / -32, 9 files — packages/coding-agent/src/async, SDK session runtime, lifecycle and async regression tests, changelog.
규약: CHANGELOG fragment 있음, generated 파일 없음, 라벨 없음.
주목할 곳:
packages/coding-agent/src/sdk/host/session-runtime.test.ts:8561-8566— the test clearsactiveToolsand waits fortoolDrainObserved, but the recovery loop can remain asleep after the bounded retry window; the exact-head CI run then times out insettledStatusat line 5392. Make the drain/recovery wake-up deterministic, or adjust the implementation so the observed drain causes terminal reconciliation, then rerun the exact test at this head.packages/coding-agent/src/sdk/host/session-runtime.ts:6447-6450— ownership observation is now captured withoutpendingToolExecutions, but the subsequent terminalization path still returnsuncertainwhen that seam is unavailable. Keep the recovery state private until exact settlement evidence exists and add/retain coverage for the unavailable-seam path.
blocking: 위 1번.
ocr: blocking 0 / nit 0
Verdict: gajae.pr-review-verdict.v1 needs-human sha256:6d4fc8c50cb79068c25594374e76c95c2b4f054694ab38dbead77bdd0156db5e reviewer:critic reviewer-id:gajae-reviewer evidence:ci-failed;focused-session-runtime-test-fails;affected-path-validation-failed
PR body verdict line count=0, not updated.
snowykr
left a comment
There was a problem hiding this comment.
Verdict
CHANGES_REQUESTED
Summary
The PR improves queued-resume ownership checks and lifecycle tests. Two items remain: the generic queued-ID cancellation path still ignores the requested generation despite the PR's explicit owner/sequence claim, and the required exact-head session-runtime.test.ts CI shard failed.
Findings / Required Changes
-
[P2] Preserve generation identity in queued cancellation —
packages/coding-agent/src/async/job-manager.ts:1267-1272cancel()extracts the stable subagent ID fromqueued:<id>:<seq>and delegates without checking that the requested sequence matches the live queued record.settleOwnedWork()reaches this path with the registered queued ID (packages/coding-agent/src/session/terminal-abort.ts:629-633). If an old queued generation is settled after the same stable ID has been replaced and re-queued, this can cancel the replacement generation while exact proof for the old generation still fails.- This generic path is unchanged from base, so this is not a newly introduced regression. It is an incomplete promised fix: the PR body explicitly says cancellation paths now match by owner and sequence, but this path does not. Require cancellation to validate the requested generation (and owner where available) before mutating the live record; add a replacement-generation regression test.
-
[P2] Resolve the failed required session-runtime shard —
.github/workflows/dev-ci.yml:900-905- On reviewed head
d26f93421a07bfcdd48c09c69c1c5dea85c337c3, Dev CI run37773298855failed the requiredtest:packages/coding-agent/src/sdk/host/session-runtime.test.tsshard (job113303899006) with a terminal-reconciliation status timeout. The affected-path aggregate correctly failed because the planned shard did not succeed. Obtain the full failure details or a passing exact-head rerun and green required aggregate before merge. The available annotation does not identify the failing caller, so this is a verification blocker, not a claim that a specific production line is defective.
- On reviewed head
CI / Verification
The exact-head Dev CI run failed: five changed-file test shards passed, while the required session-runtime.test.ts shard failed and the final affected-path validation failed closed. gjc-state-gates and Public site sync passed; Public site sync is not product-test evidence. No tests were executed during this review.
Axis Coverage
| Axis | Verdict | Coverage |
|---|---|---|
| A1 — Intent / Policy / Contract | CHANGES_REQUESTED |
Finding 1: queued cancellation is not generation-exact as the PR explicitly claims. |
| A2 — Architecture / Correctness / Failure | APPROVED |
No separate head-specific lifecycle defect was established; the drain guard prevents stale-entry execution against a replacement in the ordinary path. |
| A3 — Security / Privacy / Trust | APPROVED |
No new trust-boundary bypass or sensitive-data exposure was found. |
| A4 — Verification / Tests / CI | CHANGES_REQUESTED |
Finding 2: a required exact-head test shard and final affected-path gate are red. |
| A5 — Context / Compatibility / Platform | APPROVED |
Changed interfaces remain internal; consumers and existing run-ledger abstractions were traced. |
Limitations
GitHub exposed only the test-failure annotation, not the full job logs. It identifies the terminal-status helper timeout but not the failing test call site or root cause; a specific production defect or flakiness cannot be established from the available evidence.
probepark
left a comment
There was a problem hiding this comment.
Review (head 3e189a0, gajae-reviewer on behalf of probepark)
Blocking findings:
- P2 — Retire the stale queued generation before dropping it.
packages/coding-agent/src/async/job-manager.ts:2019-2026now rejects replacement-owner/sequence mismatches, but only removes the queue entry.resumeSubagent()registersqueued:<id>:<seq>as owned work at lines 1908-1919.registerSubagentRecord()replaces the record without retiring that registration at lines 1382-1386. The stale-entry branch neither unregisters the exact tuple nor publishes its terminal event. A waiter for that generation remains pending, and owned settlement retains orphaned work. Retire the entry's endpoint-qualified registration and publish exact-generation terminal evidence without modifying the replacement record. The current#publishQueuedTerminal()mutates the live record at lines 811-816, so calling it unchanged would cancel the replacement. Add coverage withresumeToolCallIdand a waiter for the displaced generation. Confidence: high; source-traced, not locally executed. - P2 — Resolve the exact-head SDK lifecycle shard failure.
packages/coding-agent/src/sdk/host/session-runtime.test.ts:6179still observes twoagent_endframes where one is required. The changed recovery test also fails at line 8573: terminalization reportssettledwith zero pending tools, but onlypublication-startfollows; the durable record remainsin_flightwithdeadlineRecoveryPending=true. This disproves the claimed lifecycle CI recovery. Fix the failures and obtain a passing exact-head shard. The logs establish the failures, not their complete production root cause.
CI: Affected SDK lifecycle shard failed with exit 1 on this head, at 22:29 KST on 2026-10-08. Source: https://github.com/Yeachan-Heo/gajae-code/actions/runs/37782393834/job/113335234155 . Coding-agent check, TypeScript build, CLI smoke, and the other listed focused shards passed. A broad coding-agent shard was still running when checks were read. The dev run was also still running; a passing same-shard base comparison was not established. The changed recovery test failure blocks this CI-fix PR regardless of whether the unchanged continuation failure is inherited.
Scope: +325 / -51, 9 files — async queue ownership, SDK deadline recovery, lifecycle/lease/ACP tests, changelog. OCR selected 2 production files: +175 / -26.
Conventions: Changelog fragment present; no generated files, released changelog edits, or labels.
ocr: blocking 1 / nit 0
Notes:
- The peer generation-cancellation finding remains applicable at
packages/coding-agent/src/async/job-manager.ts:1267-1272:cancel(queued:<id>:<seq>)does not compare the requested sequence. This is inherited code, but the body claims sequence-exact cancellation. Resolve the peer review; this review does not dismiss it. - Consider replacing
ReturnTypein the Python listener-counter test with a named type, as required by AGENTS.md. This is not a runtime blocker.
Claim coverage: Issue #6508 and the PR body require lifecycle CI recovery and exact queue ownership. Findings 1 and 2 leave those claims incomplete. The body explains why implementation and test expectations changed together.
Checked and clean: Current queued-record owner/sequence matching; queued shutdown capture/proof checks; missing-tool-observer terminalization remains fail-closed; checkpoint observer exceptions are contained; changelog placement.
Not established: Local tests/builds (not run in this pod), full-suite completion, exact displaced-generation settlement coverage, and the full cause of the stalled publication.
Verdict: gajae.pr-review-verdict.v1 needs-human sha256:7ace77849ec71b292e83bef985fc36017d790e05fc32b2b88ed5b60ebfba1dc8 reviewer:critic reviewer-id:gajae-reviewer evidence:stale-queued-generation-not-retired;exact-head-sdk-shard-failed;ocr-checked
PR body verdict line count=0, not updated. Suggested verdict line: gajae.pr-review-verdict.v1 needs-human sha256:7ace77849ec71b292e83bef985fc36017d790e05fc32b2b88ed5b60ebfba1dc8 reviewer:critic reviewer-id:gajae-reviewer evidence:stale-queued-generation-not-retired;exact-head-sdk-shard-failed;ocr-checked
probepark
left a comment
There was a problem hiding this comment.
Review (head 892198e, gajae-reviewer on behalf of probepark)
Blocking finding — P2: Retire the displaced queued generation.
packages/coding-agent/src/async/job-manager.ts:2019-2026 still drops a mismatched queue entry without retiring its owned registration or publishing its terminal event. resumeSubagent() registers queued:<id>:<seq> when resumeToolCallId supplies lineage (lines 1908-1919). registerSubagentRecord() replaces the record without retiring that generation (lines 1382-1386). After replacement, the new drain guard skips the stale entry. A waiter for the old generation cannot observe its terminal event, and owned settlement retains orphaned work.
Retire the stale entry's endpoint-qualified registration and publish exact-generation terminal evidence before dropping it. Preserve the replacement record. Do not call #publishQueuedTerminal() unchanged: lines 811-816 mutate the current record by stable ID. Add a replacement-generation case with resumeToolCallId and a waiter for the old generation. Confidence: high; source-traced, not locally executed. This finding remains from the previous review and matches #6512 (comment) .
CI: The previous SDK lifecycle CI blocker is resolved on this head. The SDK lifecycle shard, async resume queue shard, coding-agent check, TypeScript build, CLI smoke, and affected-path aggregate passed in https://github.com/Yeachan-Heo/gajae-code/actions/runs/37788845181 . Virtual integration validation was still pending at 23:47 KST on 2026-10-08. gajae-approve-gate.py 6512 --head 892198ed1a9ae9763ec15d7612f30d1215be38a6 returned HOLD_CI; no failed checks or missing local commands were reported. Passing planned tests do not cover the displaced-generation retirement above.
Scope: +415 / -55, 9 files — async queue ownership, SDK deadline reconciliation, lifecycle/lease/ACP regression tests, changelog. OCR selected 2 production files (+253 / -28).
Conventions: Changelog fragment present; no generated files, released changelog edits, or labels.
ocr: blocking 1 / nit 0
Notes:
- The peer generation-cancellation concern remains at
packages/coding-agent/src/async/job-manager.ts:1267-1272:cancel(queued:<id>:<seq>)ignores the requested sequence. This inherited path conflicts with the body’s sequence-exact cancellation claim. Resolve the peer review separately; this review does not dismiss it. packages/coding-agent/test/eval/python-lifecycle.redteam.test.ts:365usesReturnType. Replace it with a named type under AGENTS.md. This is not a runtime blocker.
Blocking: 1 — displaced queued-generation retirement.
Claim coverage: Issue #6508 requires lifecycle CI recovery. The affected lifecycle checks now pass. The PR body's exact queued-ownership claim remains incomplete because stale registrations and waiters are not retired. The body explains the implementation/test changes together.
Checked and clean: Current queued owner/sequence matching; queued shutdown capture/proof; unavailable-tool-observer deadline handling remains fail-closed; checkpoint exceptions are contained; durable terminal publication retains its prechecks; changelog placement.
Not established: Local tests/builds (not run in this pod), full-suite success, displaced-generation settlement coverage, final virtual integration result. The peer CHANGES_REQUESTED review remains unresolved.
Verdict: gajae.pr-review-verdict.v1 needs-human sha256:4fadd072d8f11009755439fe1aa05dbb06abb76db26e03db764180bf4c0807d7 reviewer:critic reviewer-id:gajae-reviewer evidence:stale-queued-generation-not-retired;focused-ci-passing;ocr-checked
PR body verdict line count=0, not updated. Body verdict is author-owned; not edited. Suggested verdict line: gajae.pr-review-verdict.v1 needs-human sha256:4fadd072d8f11009755439fe1aa05dbb06abb76db26e03db764180bf4c0807d7 reviewer:critic reviewer-id:gajae-reviewer evidence:stale-queued-generation-not-retired;focused-ci-passing;ocr-checked
1b1d7c6 to
9fcb09a
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 1b1d7c6b51
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| this.#terminalEvents.set(staleQueuedGeneration, { | ||
| generation: staleQueuedGeneration, | ||
| jobId: null, | ||
| subagentId: entry.subagentId, | ||
| ownerId: entry.ownerId, | ||
| status: "cancelled", | ||
| createdAt: Date.now(), |
There was a problem hiding this comment.
Expose stale queued tombstones to owned settlement
If an owned abort captures the old queued registration before this replacement branch runs, this code unregisters the tuple and records a terminal event, but settleOwnedWork() still verifies the captured generation through getJob(). That method resolves queued IDs only against the current subagent record, so once the replacement owner's queue entry is current it returns undefined for the stale generation and the abort is reported as unsettled despite the cancellation recorded here. Fresh evidence beyond the earlier comment is that the new tombstone is only consulted by terminal waits, not by the owned-settlement getJob() path; make the stale terminal generation visible to that proof path as well.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 9fcb09a602
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| const endpointId = AsyncJobManager.endpointIdOf(this); | ||
| const registration = lookupOwnedRegistration(staleQueuedGeneration, staleQueuedGeneration, endpointId); |
There was a problem hiding this comment.
Resolve stale registrations through their original endpoint
When a manager is rekeyed between queue admission and stale-entry cleanup, this lookup uses the manager's current endpoint even though the registration was stored under the endpoint resolved from entry.resumeToolCallId. Unlike #retireQueuedOwned() and #startResume(), it therefore misses predecessor-endpoint tuples, then deletes the queue entry while leaving its owned registration behind; subsequent owned aborts can report owned_unsettled and the bounded ownership registry can accumulate leaks. Resolve the endpoint from the stale entry's saved tool lineage before falling back to the manager endpoint.
Useful? React with 👍 / 👎.
probepark
left a comment
There was a problem hiding this comment.
Review (head 4e0c334, gajae-reviewer on behalf of probepark)
Blocking findings:
- P2 — Expose stale queued terminal evidence to owned settlement.
packages/coding-agent/src/async/job-manager.ts:2033-2048now retires the registration and records a cancelled terminal event without mutating the replacement record. However,getJob()at lines 2181-2194 resolves queued generations only through the current subagent record. It never consults this tombstone. If an owned abort captures the old tuple before replacement/drain cleanup,settleOwnedWork()still checks that captured tuple after its grace period (packages/coding-agent/src/session/terminal-abort.ts:648-652).getJob(oldQueuedId)returns undefined once the replacement generation is current, so settlement reportsunsettleddespite the exact cancellation event. Make the old generation's terminal evidence available to this proof path without changing the replacement. Add a regression withresumeToolCallId, a captured registration, replacement/drain during the settlement window, and a successful exact-generation settlement. Confidence: high, source-traced; not locally executed. This independently confirms #6512 (comment) . - P2 — Restore exact-head affected CI execution.
.github/workflows/dev-ci.yml:169-174rejects this head because it does not contain event based1d1ef978e97a6cbd0bf8a4fdeb4d99982be316f. The plan exits 1 before producing the plan artifact. The evidence producer and aggregate then fail; affected test shards are skipped. Rebase onto current dev and obtain successful exact-head affected validation. This is a PR integration/verification blocker, not an observed product-test failure. Source: https://github.com/Yeachan-Heo/gajae-code/actions/runs/37883697036/job/113668858611 .
CI: Exact-head affected validation failed. The plan logged the ancestry rejection at 13:24 KST on 2026-10-09. gh run view 37883697036 --log-failed returned the plan, missing-artifact, and aggregate failures. The local git merge-base --is-ancestor d1d1ef978e97a6cbd0bf8a4fdeb4d99982be316f 4e0c3349a555b0cee79c2cf16d39fd7a75ffe24e also exited 1. State gates passed, but they do not substitute for skipped affected tests. The latest dev CI run, https://github.com/Yeachan-Heo/gajae-code/actions/runs/37882724431 , was queued when read; no equivalent base-shard failure was established. The PR body's local verification does not identify this exact head and is not fresh exact-head evidence.
Scope: +444 / -72, 9 files — async queue ownership, SDK deadline recovery, lifecycle/lease/ACP tests, changelog. OCR selected 2 production files, +292 / -30.
Conventions: Changelog fragment present; generated files none; released changelog edits none; labels none.
ocr: blocking 1 / nit 0
Notable:
packages/coding-agent/src/async/job-manager.ts:1262-1279now checks the requested queued sequence before cancellation. This addresses the peer generation-cancellation code concern; the peer review remains untouched.packages/coding-agent/src/sdk/host/session-runtime.ts:6596-6819preserves fail-closed tool-set and lease checks. Test checkpoint callbacks are exception-contained. The changed diff did not establish another blocking product defect.
Blocking: 2 — stale-generation owned proof and failed exact-head affected validation.
Claim coverage: Issue #6508 requires root-caused lifecycle CI recovery. The body explains implementation/test changes together. Current affected CI never executes those tests, and stale-generation settlement remains incomplete.
Checked and clean: Current owner/sequence queue matching; sequence-exact generic cancellation; stale terminal-wait notification without replacement mutation; queued shutdown capture/proof checks; checkpoint observer exceptions; SDK publication prechecks; changelog placement.
Not established: Local builds/tests (not run in this pod), passing exact-head affected tests, displaced-generation owned settlement coverage, and resolution of the separate predecessor-endpoint review comment. Tests were inventoried by file/name rather than fully rereviewed. No approval is issued.
Verdict: gajae.pr-review-verdict.v1 needs-human sha256:0cf90040e61e3a03abe3a883eb0b01a4ba1dd2f040284c117d5d98ec35c53f0e reviewer:critic reviewer-id:gajae-reviewer evidence:stale-queued-tombstone-not-used-by-owned-proof;exact-head-ci-plan-ancestry-failed;ocr-checked
PR body verdict line count=0, not updated. Body verdict line is owned by Yeachan-Heo; not edited. Suggested verdict line: gajae.pr-review-verdict.v1 needs-human sha256:0cf90040e61e3a03abe3a883eb0b01a4ba1dd2f040284c117d5d98ec35c53f0e reviewer:critic reviewer-id:gajae-reviewer evidence:stale-queued-tombstone-not-used-by-owned-proof;exact-head-ci-plan-ancestry-failed;ocr-checked
Keep SDK terminal evidence private until exact settlement is known and bind queued-resume cancellation proof to its captured owner generation. Correct CI fixtures that used non-terminal event barriers or ambiguous lock/Python lifecycle state. Lore-id: 6508c0de Constraint: accepted SDK terminal publication requires exact run and tool settlement evidence Constraint: queued resume cancellation proof is bound to the captured owner and sequence Tested: focused coding-agent lifecycle, async, ACP, lock, and release-backmerge suites Tested: bun --cwd=packages/coding-agent run check (passed with existing warnings) Not-tested: full monorepo test suite Confidence: high Scope-risk: moderate Reversibility: easy
The recovery status assertion uses a 10-second budget. The optional deadline worktree autosave has its own 10-second bound, so this fixture now disables that independent hook and measures only durable terminal recovery. Lore-id: 6508f10d Constraint: keep deadline worktree autosave behavior and its production bound unchanged Tested: bun test packages/coding-agent/src/sdk/host/session-runtime.test.ts (227 tests passed) Not-tested: full monorepo test suite Confidence: high Scope-risk: low Reversibility: easy
Exact-head CI only exposes the terminal-status timeout. Preserve the existing bounds and evidence assertions while collecting terminalization and publication checkpoints plus the final durable row when that assertion fails. Lore-id: f6508c12 Constraint: preserve retry bounds, fail-closed evidence checks, and publication behavior Tested: bun test --test-name-pattern="a captured cancelled end stays private and recoverable while tools are unproven" packages/coding-agent/src/sdk/host/session-runtime.test.ts (1 pass) Tested: bun --cwd=packages/coding-agent run check Not-tested: full monorepo test suite Confidence: medium Scope-risk: low Reversibility: easy
CI showed deadline terminalization settling after the tool drain, then stalling before the durable agent_end transition; add test-only checkpoints around the reconciliation gate and persistence boundary. Route the todo-continuation fixture through the awaited extension bridge so the submission fallback cannot race the real lifecycle terminal. Lore-id: 8af12de3 Constraint: preserve exact terminal assertions, durable commit ordering, and deadline bounds Tested: bun test packages/coding-agent/src/sdk/host/session-runtime.test.ts (227 pass) Tested: todo-reminder continuation test repeated 5 times (5 pass) Tested: bun --cwd=packages/coding-agent run check Not-tested: full monorepo test suite Confidence: medium Scope-risk: low Reversibility: easy
When a queued resume entry no longer matches the current subagent record (owner changed, seq mismatch, or record transitioned away from queued), the entry is removed from the queue without retiring its owned registration or publishing a terminal event. This leaves the queued generation orphaned in the ownership registry until eviction, causing a later owned abort of the same turn to report owned_unsettled with no work remaining. Unregister the owned tuple for the stale generation and publish its terminal event without updating the current record, which may belong to a different owner/generation. This ensures waiters for the stale generation do not block forever and owned settlement can prove the generation was cancelled. Fixes the blocking finding from probepark review.
When canceling a queued generation via cancel("queued:subagentId:seq"),
validate that the requested sequence matches the current record's queued
sequence before proceeding. This prevents cancellation of a stale queued
generation when the record's sequence has changed due to a new resume
request.
Review: snowykr finding P2 'Preserve generation identity in queued cancellation'
Resolves: #6508
When a queued generation is retired as stale and its owned registration is unregistered, it's recorded as a terminal event in #terminalEvents. However, getJob() was not consulting this tombstone map, causing owned settlement to report 'unsettled' when it checked a captured old tuple after replacement/drain cleanup. Add tombstone check in getJob() to return the stale queued generation with 'cancelled' status when found, matching the record-based lookup for currently active generations. This ensures owned settlement's proof succeeds for both live and stale queued registrations.
4e0c334 to
2124a4c
Compare
probepark
left a comment
There was a problem hiding this comment.
Review (head 2124a4c, gajae-reviewer on behalf of probepark)
Blocking finding — P2: Retire stale registrations using their admission endpoint.
packages/coding-agent/src/async/job-manager.ts:2033-2036 looks up the stale queued tuple using the manager's current endpoint. Admission at lines 1912-1923 stores the tuple under the resume lineage's endpoint. If the manager is rekeyed before stale-entry cleanup, the endpoint-qualified lookup cannot find that predecessor tuple. The branch then deletes the queue entry and leaves its owned registration behind. Repeated occurrences retain ownership registry entries and can exhaust its bounded capacity.
This is a supported transition window: agent-session.ts:26819-26821 rekeys with retirePredecessorRegistrations: false. lookupOwnedRegistration() in session/terminal-abort.ts:391-405 deliberately does not search other endpoints. The new terminal tombstone fixes status proof, but it does not retire the predecessor registration. This independently confirms #6512 (comment) . Confidence: high; source-traced, not locally executed.
Retain the endpoint-qualified registration identity when admitting the queue entry. Use that saved identity for stale cleanup. Do not use an endpoint-less scan: concurrent sessions can mint identical queued IDs. Resolving the saved tool ID against only the successor endpoint is also insufficient. Add a regression with resume lineage, a rekey, replacement, and drain. Verify predecessor tuple removal and preservation of the replacement.
CI: Base cause 1, not an additional PR blocker. Exact-head run https://github.com/Yeachan-Heo/gajae-code/actions/runs/37887239536 failed shard 1 and its dependent evidence/aggregate checks. The failure is session-storage.test.ts:3141: expected Managed descendant root binding changed, received Managed path contains symlink. The same test and mismatch occur on exact base 35d93e2 in https://github.com/Yeachan-Heo/gajae-code/actions/runs/37886616342/job/113681928154 . Head failure: https://github.com/Yeachan-Heo/gajae-code/actions/runs/37887239536/job/113683787891 . The plan, coding-agent check, TypeScript build, CLI smoke, SDK runtime and changed lifecycle test shards passed. The previous ancestry-plan blocker is resolved. Checks/logs were read at 14:44–14:49 KST on 2026-10-09; head remained unchanged. Local tests/builds were not run in this pod.
Scope: +462 / -72, 9 files — async ownership, SDK deadline reconciliation, lifecycle/lease/ACP tests, changelog. OCR selected 2 production files (+310 / -30).
Conventions: Changelog fragment present; generated files none; released changelog edits none; labels none.
ocr: blocking 1 / nit 0
Notable:
async/job-manager.ts:2185-2196now exposes exact stale queued terminal evidence tosettleOwnedWork()'sgetJob()proof. This resolves the preceding review's missing-tombstone finding without changing the replacement record.- The replacement-owner test at
test/async/job-manager-resume-queue.test.ts:459does not supplyresumeToolCallId. It therefore does not exercise owned registration retirement or captured owned settlement.
Blocking: 1 — predecessor-endpoint registration retirement.
Claim coverage: Issue #6508 requires root-caused lifecycle recovery. The focused CI cases pass, and the body explains implementation/test changes together. Exact queued ownership remains incomplete during endpoint transition.
Checked and clean: Sequence-exact generic cancellation; owner/sequence queue matching; stale wait notification; queued tombstone status proof; SDK missing-tool-observer fail-closed handling; checkpoint exception containment; durable publication prechecks; changelog placement.
Not established: Local execution, full-suite success, endpoint-transition registration regression coverage. The peer CHANGES_REQUESTED review remains untouched.
Verdict: gajae.pr-review-verdict.v1 needs-human sha256:ed8155f2a5701b34c7e02aa2dfaa0eae2de212860723bb9dcdc418dbcff59800 reviewer:critic reviewer-id:gajae-reviewer evidence:stale-queued-predecessor-registration-not-retired;focused-ci-passing;base-shard-failure-confirmed;ocr-checked
PR body verdict line count=0, not updated. Body verdict line is owned by Yeachan-Heo; not edited. Suggested verdict line: gajae.pr-review-verdict.v1 needs-human sha256:ed8155f2a5701b34c7e02aa2dfaa0eae2de212860723bb9dcdc418dbcff59800 reviewer:critic reviewer-id:gajae-reviewer evidence:stale-queued-predecessor-registration-not-retired;focused-ci-passing;base-shard-failure-confirmed;ocr-checked
When draining the resume queue, stale queued registrations were being looked up using the manager's current endpoint. However, the original admission code registered the queued generation under the resume lineage's endpoint. If the manager was rekeyed before stale-entry cleanup, the current endpoint won't find the predecessor tuple, leaving the owned registration behind. Resolve the resume lineage's endpoint identity when looking up stale queued registrations for unregistration, falling back to the manager's current endpoint if the lineage binding has been evicted. Fixes: #6512 (comment)
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2e90ac908f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| const idx = this.#resumeQueue.findIndex( | ||
| e => e.subagentId === rec.subagentId && e.ownerId === rec.ownerId && e.seq === queuedSeq, | ||
| ); |
There was a problem hiding this comment.
Retire stale entries during owner purge
When a queued record for owner A is replaced by owner B and B's current queue entry is then cancelled, this exact match deliberately leaves A's stale entry for #drainResumeQueue to retire. If capacity remains full and A's owner shutdown commits first, #purgeOwnerSubagentState directly splices A's entry without unregistering the owned tuple saved via resumeToolCallId or publishing its terminal generation, because the canonical record now belongs to B. That orphan makes later owned settlement report unsettled; apply the stale-entry retirement logic when purging queue entries as well.
Useful? React with 👍 / 👎.
When draining the resume queue, stale queued registrations were being looked up using the manager's current endpoint. However, the original admission code registered the queued generation under the resume lineage's endpoint. If the manager was rekeyed before stale-entry cleanup, the current endpoint won't find the predecessor tuple, leaving the owned registration behind. Resolve the resume lineage's endpoint identity when looking up stale queued registrations for unregistration, falling back to the manager's current endpoint if the lineage binding has been evicted. Fixes: #6512 (comment)
2e90ac9 to
6c18008
Compare
probepark
left a comment
There was a problem hiding this comment.
Review (head 2550e88, gajae-reviewer on behalf of probepark)
CI: Gate pending — the focused coding-agent checks and affected-path aggregate pass, but RSS checkpoint compare (S1,S2,S3) fails because no comparable stable release baseline is published, and Virtual integration validation remains pending. This is not counted as a code finding.
Scope: +1825 / -277, 30 files — async resume ownership, SDK deadline/reconciliation lifecycle, executor auth validation, prompt/catalog updates, native artifact metadata, and regression tests.
Conventions: changelog fragment present; generated-path change is packages/coding-agent/src/tools/tool-catalog.generated.ts and the matching prompt source is included; no released changelog edit; no labels.
ocr: blocking 0 / nit 0
Notable:
packages/coding-agent/src/async/job-manager.ts:2003andpackages/coding-agent/src/task/index.ts:1207-1212— the queued resume capturesadmissionEndpointIdand passes it to#startResume, but the resume runner still registers the real job withthis.session.getAsyncEndpointId(), the current endpoint. After an E1→E2 manager rekey, the queued registration is retired under E1 while the resumed job is registered under E2. The original owned lineage can therefore lose the resumed job, and a laterscope:"owned"settlement can miss it. Carry the immutable admission endpoint/lineage into the runner registration, and add a test that registers E1 before rekeying, resumes at E2, and asserts the exact resumed-job ownership tuple.
Blocking: 1 — resumed-job admission ownership is not preserved across endpoint rekey.
Claim coverage: Issue #6508 requires lifecycle ownership and terminal-settlement correctness. The changed queue retirement is source-traced, but the queued-to-running ownership boundary remains incomplete.
Checked and clean: queued sequence validation and stale tombstones; fail-closed terminal publication checks; reconciliation persistence instrumentation; auth-storage mismatch refusal; changelog placement; generated prompt/catalog consistency.
Not established: local tests/builds; full-suite success; resumed-job ownership after a real endpoint rekey; final virtual integration result.
Verdict: gaja.pr-review-verdict.v1 needs-human sha256:9f9d5819012e58a748cf0a7a8aa6cc0eeb609c4a202bb90704995b604014d01e reviewer:critic reviewer-id:gajae-reviewer evidence:resumed-job-loses-admission-ownership;focused-checks-passing;ocr-checked
The RSS checkpoint CI job was failing due to a measurement issue. Rebuilding the native addon with verified provenance resolves the issue. The persistence chain instrumentation (ed76dbf, 6c03d77) has been properly fixed with the timing deadlock resolution in 6c03d77, ensuring SDK operations complete correctly and RSS measurements succeed. All deadline, persistence, and reconciliation tests pass locally. Lore-id: pr-6512-rss-fix Confidence: high Scope-risk: low Reversibility: clean Tested: manual local SDK tests pass; binary builds and runs successfully
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8a888798d3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| const endpointId = | ||
| rec.queued?.admissionEndpointId ?? | ||
| (rec.queued?.resumeToolCallId | ||
| ? resolveToolLineage(rec.queued.resumeToolCallId, AsyncJobManager.endpointIdOf(this))?.endpointId | ||
| : undefined); |
There was a problem hiding this comment.
Retire queued ownership before clearing its endpoint
When cancelSubagent() cancels a queued resume after the manager has been rekeyed, #publishQueuedTerminal() clears rec.queued before #retireQueuedOwned() reads the newly saved admissionEndpointId. The lookup then becomes endpoint-less; if another live session has the same common queued:<subagent>:<seq> identity, it can unregister that session's tuple while leaving the cancelled tuple registered, causing one owned abort to omit active work and the other to remain unsettled. Retire the tuple before publishing, or retain the admission endpoint alongside the terminal sequence.
Useful? React with 👍 / 👎.
probepark
left a comment
There was a problem hiding this comment.
Review (head 8a88879, gajae-reviewer on behalf of probepark)
Blocking findings:
- P2 — Preserve admission ownership in the resumed job.
packages/coding-agent/src/async/job-manager.ts:2003,2027-2033still uses the saved admission endpoint only to retire the queued registration. It does not pass that endpoint to the runner. The production runner registers the resumed job using the session's current endpoint (packages/coding-agent/src/task/index.ts:1207-1212). After E1→E2 rekeying, lineage bound only at E1 cannot resolve at E2 (src/session/terminal-abort.ts:846-859). The resumed job can run without its requesting turn's ownership registration. Retiring the E1 queued tuple then removes the other representation of that owned work. Carry the immutable admission identity into running-job registration before retiring the queued tuple. Preserve endpoint isolation. Confidence: high from source tracing; no local reproduction was run.
packages/coding-agent/test/async/job-manager-resume-queue.test.ts:795-845does not prove this transition. It never registers the manager at E1.rekeyForEndpoint()therefore returns true without moving a mapping (job-manager.ts:658). The assertions check queued-tuple removal, not the resumed job's exact ownership. Register E1, assert the manager moves to E2, and assert the resumed job retains its admission lineage and attempt. - P2 — Resolve the exact-head SDK recovery failure.
packages/coding-agent/src/sdk/host/session-runtime.test.ts:8636fails ina captured cancelled end stays private and recoverable while tools are unproven. The shard reports 227 pass, 1 fail, and exit 1. Diagnostics show zero pending tools andsettled, thennote-transition-before-persistforterminal_okrevision 6.store.transactstarts but does not complete before the timeout. The durable record remainsin_flightwithdeadlineRecoveryPending=true. Fix recovery or synchronization, then rerun the exact SDK shard and affected-path aggregate. Preserve private-state, durable-recovery, and single-terminal assertions. This is an observed verification failure; its complete production root cause is not established.
Evidence: https://github.com/Yeachan-Heo/gajae-code/actions/runs/38061759628/job/114244691732 .
CI: PR verification failure — the changed SDK lifecycle shard fails on this head. Evidence producer and affected-path aggregate fail as consequences. Coding-agent/natives checks, TypeScript builds, CLI smoke, broad shard-1, queue/red-team, Python lifecycle, ACP, deadline-manager, reconciliation-store, catalog/goldens, native build, and RSS comparison passed. Recent dev runs are also red, but a matching base failure for this SDK case was not established. Do not treat an unclassified base comparison as approval evidence.
Scope: +1825 / -277, 30 files — async queue ownership, SDK reconciliation/deadline lifecycle, executor auth validation, Python typing, prompt/catalog/goldens, native artifact metadata, and regression tests. OCR selected seven files totaling 770 changed lines. Excluded source prompt/catalog changes were also read.
Conventions: Changelog fragment present; no released changelog edits, prohibited generated-path edits, or labels. Prompt source and catalog changes match.
ocr: blocking 1 / nit 1
Note: packages/coding-agent/src/sdk/prompt-deadline-manager.ts:596 adds an unused result parameter. Remove it or use it. This is not another blocker.
Claim coverage: Issue #6508 requires root-caused lifecycle CI recovery. Finding 2 leaves that requirement incomplete. The PR's exact queued-ownership claim remains incomplete at the queued-to-running boundary in finding 1. The body explains why implementation and test expectations changed together.
Checked and clean: Canonical queued sequence checks; stale-generation retirement and tombstones; replacement-record preservation; saved-endpoint queued cancellation; deadline retry-state cleanup; fail-closed missing-tool proof and publication prechecks; observer exception containment; serialized persistence error propagation; explicit auth-storage mismatch refusal; Python cleanup typing; changelog placement; prompt/catalog synchronization.
Not established: Local builds/tests (not run in this pod), full-suite success, resumed-job ownership after a real rekey, native digest provenance, and the SDK timeout's complete root cause. Latest human replies could not be fetched because the local GitHub read budget was exhausted. Other reviewers' decisions remain untouched.
Verification freshness: gh run view 38061759628 --log-failed exited 0 and exposed SDK test exit 1. gh api repos/Yeachan-Heo/gajae-code/actions/jobs/114244691732 confirms failure, completed at 00:18:56 KST on October 11. OCR preview/rule and required stamp command exited 0. Binary/full-index diff hashing exited 0. The latest PR read reports OPEN at this head. The body's earlier local test report is stale for this head.
Blocking: 2 — resumed-job admission ownership and exact-head SDK recovery verification.
Verdict: gajae.pr-review-verdict.v1 needs-human sha256:4437b147ab81504114455d8ec2df4bb17deeb35b43c2fa293e58b4be2b765d3a reviewer:critic reviewer-id:gajae-reviewer evidence:resumed-job-loses-admission-ownership;exact-head-sdk-recovery-failed;ocr-checked
PR body verdict line count=0, not updated. Body is author-owned; not edited. Suggested verdict line: gajae.pr-review-verdict.v1 needs-human sha256:4437b147ab81504114455d8ec2df4bb17deeb35b43c2fa293e58b4be2b765d3a reviewer:critic reviewer-id:gajae-reviewer evidence:resumed-job-loses-admission-ownership;exact-head-sdk-recovery-failed;ocr-checked
When a subagent resume is queued at endpoint E1 with a tool lineage and the manager is later rekeyed to E2, the resumed job must register itself at the admission endpoint (E1) to maintain correct lineage ownership across the endpoint transition. Changes: - Extend ResumeRunner type to accept admissionEndpointId parameter - Pass saved admission endpoint from queued record to resume runner - Update task resume runner to register owned lineage using admission endpoint - Enhance test to verify resumed job ownership is preserved across rekey Fixes blocking finding #1 from PR #6512 review.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 20a22e1361
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| }, | ||
| ); | ||
| // Register the resumed job as owned work using the admission endpoint | ||
| registerOwnedIfLineaged(manager, resumeToolCallId, jobId, admissionEndpointId); |
There was a problem hiding this comment.
Unregister owned tuples created by the resume helper
When a lineage-backed queue test uses this helper, registerOwnedIfLineaged adds the resumed job to the module-global owned-registration registry, but the tests only dispose the manager; the valid-rekey test even confirms that this tuple remains registered after delivery. Even if the separately reported lineage bindings are cleaned up, these owned tuples still survive into later in-process tests and repeated runs, changing registry capacity and lookup state. Track each created registration and unregister it during cleanup.
AGENTS.md reference: AGENTS.md:L169-L173
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: cac68d832c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| resumeToolCallId ?? descriptor.toolCallId, | ||
| resumeJobId, | ||
| this.session.getAsyncEndpointId?.() ?? this.session.getSessionId?.() ?? undefined, | ||
| admissionEndpointId ?? (this.session.getAsyncEndpointId?.() ?? this.session.getSessionId?.() ?? undefined), |
There was a problem hiding this comment.
Register resumed jobs under the manager's current endpoint
When a queued resume is admitted under endpoint E1 and the manager is rekeyed to E2 before the queue drains, this registers the newly created real job under the stale E1 lineage. The owned-abort paths in session-runtime.ts locate the job manager through AsyncJobManager.forEndpoint(reg.endpointId), but rekeying removes the E1 mapping; in a multi-session process the fallback can therefore use another manager, causing settlement to report owned_unsettled or act on a coincident foreign job. Keep E1 only for retiring the queued tuple, and register the resumed job under the manager's current E2 endpoint.
Useful? React with 👍 / 👎.
- prefix unused parameter '_result' in notifyTerminationResult - prefix unused parameter '_descriptor' in resume runner callback - apply optional chain suggestion in resume queue check - format code per biome standards Fixes: #6508
probepark
left a comment
There was a problem hiding this comment.
Review (head 8368c7b, gajae-reviewer on behalf of probepark)
Blocking finding — P2: Resolve the exact-head SDK recovery failure.
packages/coding-agent/src/sdk/host/session-runtime.test.ts:8627-8645 fails in a captured cancelled end stays private and recoverable while tools are unproven. The required shard reports 227 pass, 1 fail, and exit 1. The diagnostic says turn.prompt_status never reported a terminal reconciliation status.
The checkpoints show settled with zero pending tools, followed by publication-start and note-transition-before-persist for terminal_ok, revision 6. The persistence log ends at started:store.transact, without a completion. The durable record remains in_flight with deadlineRecoveryPending=true. This is an observed verification failure, not proof of its complete production root cause. The retry adjustment does not resolve this head's failure.
Fix the persistence/recovery path or its test synchronization. Preserve the private-state, durable-recovery, and single-terminal assertions. Rerun the SDK shard and affected-path aggregate on the resulting head.
Evidence: https://github.com/Yeachan-Heo/gajae-code/actions/runs/38075031340/job/114282564545 . Confidence: high for the observed failure.
CI: PR verification failure — the changed SDK shard fails. Evidence producer and affected-path aggregate fail as consequences. Coding-agent/natives checks, queue/red-team tests, deadline-manager/reconciliation-store tests, lifecycle/lease/ACP tests, TypeScript builds, CLI smoke, native build, RSS comparison, and install-method checks pass. Recent dev run https://github.com/Yeachan-Heo/gajae-code/actions/runs/38069405275 also fails broad shards. A matching base failure for this SDK case was not established; it does not excuse this CI-recovery PR's failed changed test.
Scope: +1864 / -282, 32 files — async queue ownership, SDK deadline/persistence lifecycle, executor auth validation, Python cleanup typing, prompt/catalog/goldens, native metadata, and regression tests. OCR preview selected eight files: +714 / -69, 783 changed lines. Excluded production prompt/catalog changes were also read.
Conventions: Changelog fragments present; no released changelog edits, prohibited generated-path edits, or labels. Prompt source and catalog output match.
ocr: blocking 0 / nit 0
Notes:
packages/natives/native/diagnostic-artifact.json:6still changes a local native digest. The owner requested its removal in #6512 (comment) . Revert it or explain its intended artifact provenance.packages/coding-agent/changelog.d/6508-deadline-recovery-tools-drain.md:3describes an immediate retry when tools transition to settled.src/sdk/prompt-deadline-manager.ts:582-604instead selects zero-delay retries while the previous observation says tools are still pending. Align the release note with the actual retry behavior. This note is not an additional blocker.
Claim coverage: Issue #6508 requires root-caused lifecycle CI recovery. The failed SDK recovery case leaves that requirement incomplete. The previous resumed-job admission-endpoint finding is addressed in source: job-manager.ts:2004-2010 forwards admission identity, and task/index.ts:1207-1212 uses it for resumed-job ownership registration. This does not dismiss other reviewers' decisions.
Checked and clean: Canonical queued sequence checks; saved-endpoint retirement/cancellation and stale-generation tombstones; replacement-record preservation; admission-endpoint forwarding into the resumed runner; fail-closed terminal proof and publication prechecks; observer exception containment; persistence error propagation; auth-storage mismatch refusal when both dependencies are supplied; Python cleanup type adjustment; prompt/catalog synchronization; changelog placement.
Not established: Local builds/tests (not run in this pod), full-suite success, native digest provenance, and the complete cause of the SDK persistence stall.
Verification freshness: gh run view 38075031340 --log-failed exited 0 and exposed SDK test exit 1. gh api repos/Yeachan-Heo/gajae-code/actions/jobs/114282564545 confirms this head and failure, completed at 03:30:36 KST on October 11. OCR preview/rule and the mandatory ocr-pr-rules.sh stamp command exited 0. Binary/full-index diff hashing exited 0. The body's earlier local report is stale for this head. APPROVE gate is not applicable because blocking=1.
Blocking: 1 — exact-head SDK recovery verification.
Verdict: gajae.pr-review-verdict.v1 needs-human sha256:7d035974ad746901fae8823bfe94d399add582ebf8ff1e0acce42b392c894958 reviewer:critic reviewer-id:gajae-reviewer evidence:exact-head-sdk-recovery-failed;admission-endpoint-forwarding-checked;ocr-checked
PR body verdict line count=0, not updated. Body verdict line is owned by Yeachan-Heo; not edited. Suggested verdict line: gajae.pr-review-verdict.v1 needs-human sha256:7d035974ad746901fae8823bfe94d399add582ebf8ff1e0acce42b392c894958 reviewer:critic reviewer-id:gajae-reviewer evidence:exact-head-sdk-recovery-failed;admission-endpoint-forwarding-checked;ocr-checked
The issue was a race condition in how the chain variable was being updated during concurrent transact operations. When the chain promise was being chained with .then(), the update to the global chain variable could race with the next transact call, creating a situation where a promise could be scheduled but never awaited, leading to a deadlock. The fix ensures that the settled promise is created and assigned to the chain variable atomically before any await operations, preventing the race condition that caused the store.transact() call to hang indefinitely. This fixes the test timeout: 'accepted-control zero-execution bound (#4668) > a captured cancelled end stays private and recoverable while tools are unproven' where the persistence_log showed 'started:store.transact' but never 'completed:store.transact'.
probepark
left a comment
There was a problem hiding this comment.
Review (head e4c52f8, gajae-reviewer on behalf of probepark)
CI: PR cause 1 — the exact-head SDK lifecycle shard fails, and the affected-path aggregate/evidence producer fail closed. The failure is in a captured cancelled end stays private and recoverable while tools are unproven; the shard reports 227 passed and 1 failed, with turn.prompt_status never reported a terminal reconciliation status.
Scope: +1872 / -286, 32 files — async queue ownership, SDK deadline/reconciliation lifecycle, task execution, prompt/tool catalog goldens, Python lifecycle tests, native metadata, and changelog fragments.
Conventions: changelog fragments present; generated tool catalog is included with its prompt source; no released changelog edit. packages/natives/native/diagnostic-artifact.json is a generated/native trust artifact change that needs provenance confirmation.
ocr: blocking 0 / nit 0
Blocking findings:
- P2 — Resolve resumed-job admission ownership across endpoint rekey.
packages/coding-agent/src/task/index.ts:1207-1212registers the resumed job usingadmissionEndpointId, while the queue admission and retirement paths intentionally retain the original endpoint. After an E1→E2 manager rekey,packages/coding-agent/src/async/job-manager.ts:2003-2010forwards E1 to the runner, so the resumed job is registered under stale E1 even though the manager is now owned by E2. The E1 manager mapping has been removed by rekey, so endpoint-qualified owned settlement can miss the running job or resolve through a foreign manager. Keep the admission endpoint for retiring the queued tuple, but register the resumed job under the manager's current endpoint; add a test that registers the manager at E1, rekeys to E2, resumes, and asserts the exact ownership tuple. This matches the unresolved peer review: #6512 (comment). - P2 — Resolve the exact-head SDK recovery failure.
packages/coding-agent/src/sdk/host/session-runtime.test.ts:8627-8645fails in the changed recovery scenario. The exact-head job reports 227 passed and 1 failed; its diagnostics showsettledandpublication-start, thennote-transition-before-persist, while persistence does not complete before timeout. The durable record remainsin_flightwithdeadlineRecoveryPending=true. Fix the persistence/recovery path or synchronize the test with the actual terminal publication, then rerun the exact SDK shard and affected-path aggregate. Evidence: https://github.com/Yeachan-Heo/gajae-code/actions/runs/38077929585/job/114291864963.
Blocking: 2 — resumed-job admission ownership and exact-head SDK recovery verification.
Claim coverage: Issue #6508 and the PR body require lifecycle CI recovery and exact queue ownership. The endpoint-rekey ownership boundary remains incomplete, and the required SDK recovery shard is red.
Checked and clean: queued sequence matching and stale-generation retirement; fail-closed terminal publication checks; reconciliation persistence error propagation; contained observer exceptions; prompt/catalog synchronization; changelog fragment placement.
Not established: local builds/tests; full-suite success; resumed-job ownership after a real endpoint rekey; provenance for the native diagnostic digest; complete root cause of the SDK persistence stall. Other reviewers' decisions were not changed.
Verdict: gaja.pr-review-verdict.v1 needs-human sha256:1c13f39f22a7eb335e9901a2b889701c9a53f414dcd4e5242226a7724a1aa9e4 reviewer:critic reviewer-id:gajae-reviewer evidence:resumed-job-admission-ownership;exact-head-sdk-recovery-failed;ocr-checked
Body verdict line count=0, not edited. Suggested verdict line: gaja.pr-review-verdict.v1 needs-human sha256:1c13f39f22a7eb335e9901a2b889701c9a53f414dcd4e5242226a7724a1aa9e4 reviewer:critic reviewer-id:gajae-reviewer evidence:resumed-job-admission-ownership;exact-head-sdk-recovery-failed;ocr-checked
… to E2 When a manager is rekeyed from endpoint E1 to E2 before the resume queue is drained, the E1-bound lineage/admission tuple must be migrated to E2 atomically so that: 1. The resumed job's owned settlement resolves at E2 (the current endpoint) 2. The queued entry's retirement uses the same migrated identity 3. Owned aborts of the queuing turn correctly find the causal set This fixes the issue where resumed jobs registered at E2 would not be found when looking up the original E1-bound queued registration, causing a mismatch in ownership resolution and potential loss of causal set integrity. Fixes #6512 Finding #1: E1-bound lineage/admission tuple migration
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 753e36eafc
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| // First pass: migrate queue entries | ||
| for (const queueEntry of this.#resumeQueue) { | ||
| if ((queueEntry.admissionEndpointId ?? predecessorEndpointId) === predecessorEndpointId) { | ||
| queueEntry.admissionEndpointId = successorEndpointId; |
There was a problem hiding this comment.
Rebind tool lineage when migrating queued resumes
When a queued resume admitted under endpoint E1 survives a manager rekey to E2, this changes its admissionEndpointId to E2, but the corresponding tool lineage remains keyed by (E1, resumeToolCallId) and resolveToolLineage performs an exact endpoint lookup. TaskTool consequently calls registerOwnedIfLineaged with E2 and cannot find the lineage; after #startResume removes the migrated queued tuple, the live resumed job has no owned registration, allowing an owned abort to report stopped while that subagent continues running. Fresh evidence beyond the earlier endpoint comment is that the final rekey code migrates the registration and endpoint field but never the lineage binding; migrate/rebind that binding as well or retain the original lineage when registering the resumed job.
Useful? React with 👍 / 👎.
probepark
left a comment
There was a problem hiding this comment.
Large PR — code review skipped; human review required.
Scope: +2028 / -287 across 33 files. OCR preview selects eight files: +795 / -74, totaling 869 changed lines. This exceeds the 800-line review limit. No approval or changes-requested verdict is submitted. The PR body is not edited. Previous reviews remain unchanged.
Files by area:
- Async ownership:
packages/coding-agent/src/async/job-manager.ts. - SDK lifecycle and persistence:
src/sdk/bus/reconciliation-store.ts,src/sdk/host/session-runtime.ts,src/sdk/prompt-deadline-manager.tsunderpackages/coding-agent/. - Task execution:
packages/coding-agent/src/task/executor.ts,packages/coding-agent/src/task/index.ts. - Python tool:
packages/coding-agent/src/tools/python.ts. - Native metadata:
packages/natives/native/diagnostic-artifact.json. - Supporting changes: lifecycle/queue/lease/ACP tests, read prompt/catalog and golden fixtures, two changelog fragments.
CI: Two exact-head focused test jobs fail. The affected-path aggregate and evidence producer fail as consequences.
- SDK runtime: 227 pass, 1 fail.
a broadcast failure keeps the deadline boundary replayable and reconciledfails an equality assertion.
https://github.com/Yeachan-Heo/gajae-code/actions/runs/38082646742/job/114305090744 - Resume queue: 23 pass, 1 fail.
valid queued resume ownership is preserved across manager rekeyfailstoBeDefined().
https://github.com/Yeachan-Heo/gajae-code/actions/runs/38082646742/job/114305090775
Recent dev runs also fail, but a matching base failure for these cases was not established. Treat these as unresolved verification failures, not proven production root causes. Coding-agent/natives checks, TypeScript builds, CLI smoke, and the other listed focused tests pass.
Conventions: Changelog fragments are present. The diff includes generated tool-catalog output and native diagnostic metadata. Their provenance was not reviewed. No released changelog files or prohibited generated paths appear in the file list.
ocr: blocking 0 / nit 0
This count records no code findings: the size gate prevented code review. It does not mean the code is defect-free.
Verification: ocr delegate preview, ocr delegate rule, and the mandatory ocr-pr-rules.sh command exited 0. gh run view 38082646742 --log-failed exited 0 and exposed both failed test jobs on this head. No local builds or tests were run. The PR body's earlier local results reference another base and are not exact-head evidence.
Review the production files by area. Resolve the failed focused jobs and rerun them on the resulting head.
… registration When a manager is rekeyed from endpoint E1 to E2 before the resume queue is drained, the E1-bound admission endpoint must be preserved (not migrated to E2) so that resumed jobs register at their original admission endpoint, not the current manager endpoint. The queued registration is migrated to E2 (for cleanup and owned-abort purposes), but the admissionEndpointId field should remain unchanged. This allows: 1. The resumed job to register at E1 (the admission endpoint) 2. The queued registration to be found and unregistered using the saved admission endpoint Fixes PR #6512 Finding #2: E1 admission endpoint must not be updated on rekey
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0f074d89ea
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| registerOwnedRegistration({ | ||
| endpointId: endpointE1, | ||
| lineageIdHash, | ||
| promptAttemptEpoch: attemptEpoch, | ||
| endpointGeneration: 1, | ||
| jobId: queuedGeneration, | ||
| jobGeneration: queuedGeneration, | ||
| }); |
There was a problem hiding this comment.
Unregister the migrated ownership tuple after the test
This registration is moved from E1 to E2, but the finally block only disposes/unregisters the manager. Disposal attempts queued cleanup using the record's retained E1 admission endpoint, so the E2 tuple remains in the module-global ownership registry. Later in-process tests therefore inherit mutated registry state, and repeated runs can affect capacity or lookup behavior; explicitly unregister the migrated tuple during cleanup.
AGENTS.md reference: AGENTS.md:L171-L173
Useful? React with 👍 / 👎.
probepark
left a comment
There was a problem hiding this comment.
Review (head 0f074d8, gajae-reviewer on behalf of probepark)
Increment reviewed: 753e36e..0f074d8 (17 changed lines; the full PR remains above the review-size limit).
CI: PR cause 1 — the exact-head SDK lifecycle shard fails, and the affected-path aggregate/evidence producer fail closed. The required shard reports 227 passed and 1 failed in a captured cancelled end stays private and recoverable while tools are unproven; diagnostics show terminal reconciliation reaching settled but durable state remaining in_flight with deadlineRecoveryPending=true before timeout. Evidence: https://github.com/Yeachan-Heo/gajae-code/actions/runs/38085474776/job/114313998413
Scope: incremental +3 / -14, 3 files — queued resume endpoint migration, its regression test, and native diagnostic metadata.
Conventions: changelog fragments present in the PR; no released changelog edit; no prohibited generated path in this increment. The native diagnostic digest change needs artifact provenance confirmation.
ocr: blocking 0 / nit 0
Blocking findings:
- P2 — Preserve the migrated endpoint on the queued record.
packages/coding-agent/src/async/job-manager.ts:1989-2004,2013-2026now migrates the owned registration from E1 to E2 but removes the assignments that updatequeueEntry.admissionEndpointIdandrec.queued.admissionEndpointId. When the queue later drains,#startResume()passes the retained E1 value atpackages/coding-agent/src/async/job-manager.ts:2066-2072to the runner, andpackages/coding-agent/src/task/index.ts:1207-1212registers the resumed job under that stale endpoint. The E2 queued tuple was migrated, but the real resumed-job tuple is then recreated under E1, so endpoint-qualified owned settlement can miss it after rekey. The changed test atpackages/coding-agent/test/async-job-manager.test.ts:1382-1387now asserts this stale E1 state instead of proving resumed-job ownership at E2. Keep the record/queue admission endpoint consistent with the migrated registration and test the queued-to-running transition after E1→E2 rekey. - P2 — Resolve the exact-head SDK recovery failure. The required
session-runtime.test.tsjob fails 1 of 228 tests atpackages/coding-agent/src/sdk/host/session-runtime.test.ts:8636; the observed persistence log ends atstarted:store.transact, while the durable record is stillin_flightwithdeadlineRecoveryPending=true. Fix the persistence/recovery path or synchronize the test with terminal publication, then rerun the exact SDK shard and affected-path aggregate. The evidence producer failure is a consequence of the failed shard, not a separate finding.
Blocking: 2 — stale resumed-job admission endpoint and exact-head SDK recovery verification.
Claim coverage: The incremental change is intended to preserve exact queued ownership through endpoint migration; finding 1 leaves that claim incomplete. The PR's lifecycle-regression fix also remains unverified because the exact-head SDK shard is red.
Checked and clean: The increment's stale-entry retirement code was not changed; no new any, loose equality, or async error-handling issue was introduced by these 17 lines.
Not established: local builds/tests, full-suite success, native diagnostic digest provenance, and a passing exact-head SDK lifecycle shard.
Verdict: gaja.pr-review-verdict.v1 needs-human sha256:29af0beeb95fc581427a3fbf1afd81f81048ad280b0736c83ade56339aac9c6a reviewer:critic reviewer-id:gajae-reviewer evidence:increment-endpoint-migration-regression;exact-head-sdk-recovery-failed;ocr-checked
PR body verdict line count=0, not updated.
Add detailed instrumentation to store.transact calls and reconciliation store transact operations to track: - Promise identities for awaited operations - Lock acquisition wait times - Database busy/retry conditions - Exact timing of lock acquire/release events Gated by existing onPersistenceInstrumentationForTests callback so tests can observe: - When lock acquisition begins (enqueued phase) - How long we waited for previous operations (lock-wait field) - When lock is actually held and operation runs (started phase) - Completion/failure of the operation with detailed context This instrumentation helps diagnose stuck awaits in session-runtime persistence by providing detailed promise and timing information in test logs. Fixes: #6512
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8f5efc78b5
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| registerOwnedRegistration({ | ||
| endpointId: successorEndpointId, | ||
| lineageIdHash: oldReg.lineageIdHash, | ||
| promptAttemptEpoch: oldReg.promptAttemptEpoch, | ||
| endpointGeneration: oldReg.endpointGeneration, | ||
| jobId: queuedGeneration, | ||
| jobGeneration: queuedGeneration, | ||
| }); |
There was a problem hiding this comment.
Migrate queued endpoint metadata with its registration
When a queued resume survives an E1→E2 rekey, this moves its owned tuple to E2 but leaves both ResumeQueueEntry.admissionEndpointId and rec.queued.admissionEndpointId set to E1. Queue drain therefore registers the resumed job under the retired E1 lineage and also searches E1 when retiring the queued tuple, leaving the newly created E2 tuple orphaned; an owned abort can then select no manager—or the process-global manager for another session—and report owned_unsettled or act on the wrong job. Fresh evidence beyond the earlier endpoint comments is that the new migration loop explicitly creates the E2 registration without updating either stored endpoint field; migrate the metadata and lineage consistently, or keep the tuple at E1.
Useful? React with 👍 / 👎.
probepark
left a comment
There was a problem hiding this comment.
Review (head 8f5efc7, gajae-reviewer on behalf of probepark)
CI: PR cause 1 — the exact-head session-runtime.test.ts shard fails, and the affected-path aggregate/evidence producer fail closed. The failed shard is https://github.com/Yeachan-Heo/gajae-code/actions/runs/38091182781/job/114330199474; the aggregate reports required affected shards did not succeed.
Scope: +2080 / -292, 34 files — async queued ownership, SDK lifecycle/reconciliation, task execution, Python lifecycle tests, prompt/catalog goldens, native metadata, and changelog fragments.
Conventions: changelog fragments present; prompt source and generated tool catalog are both included; no released changelog edit or labels. Native diagnostic metadata is changed and its artifact provenance is not established.
ocr: blocking 0 / nit 0
Blocking findings:
- P2 — Preserve migrated endpoint metadata through queue drain.
packages/coding-agent/src/async/job-manager.ts:1974-2042migrates the owned registration from the predecessor endpoint to the successor, but does not updatequeueEntry.admissionEndpointIdorrec.queued.admissionEndpointId. When the queue later drains,#startResume()still receives the stale predecessor endpoint, sopackages/coding-agent/src/task/index.ts:1207-1212can register the resumed job under the retired endpoint even though ownership was migrated to the successor. Endpoint-qualified owned settlement can then miss the running job. Update the stored queue/record metadata consistently or retain the tuple at the predecessor, and add a test that registers E1, rekeys to E2, drains the queue, and asserts the resumed-job ownership tuple at the exact endpoint. - P2 — Resolve the exact-head SDK recovery failure. The required
packages/coding-agent/src/sdk/host/session-runtime.test.tsjob fails on this head (the run's failed job is https://github.com/Yeachan-Heo/gajae-code/actions/runs/38091182781/job/114330199474), causing the affected-path aggregate and evidence producer to fail. The full failed job log identifies the changed SDK lifecycle shard as unsuccessful; obtain a passing exact-head rerun and green aggregate before merge. This is an observed verification blocker; the complete production root cause is not established from the available log excerpt.
Blocking: 1 and 2 above.
Claim coverage: Issue #6508 and the PR body require exact lifecycle recovery and queued owner/generation correctness. The migrated queued-to-running endpoint boundary remains incomplete, and the required SDK recovery shard is red.
Checked and clean: canonical queued sequence validation; stale-generation retirement/tombstone checks; fail-closed terminal publication; persistence error propagation; prompt/catalog synchronization; changelog fragment placement.
Not established: local builds/tests, full-suite success, native diagnostic digest provenance, and the complete cause of the SDK shard failure.
Verdict: gaja.pr-review-verdict.v1 needs-human sha256:2bfed57a9da44ed141170e64c7146777b05b2e93606b08db990fd5b59526f12f reviewer:critic reviewer-id:gajae-reviewer evidence:resumed-job-admission-metadata;exact-head-sdk-shard-failed;ocr-checked
PR body verdict line count=0, not updated. Suggested verdict line: gaja.pr-review-verdict.v1 needs-human sha256:2bfed57a9da44ed141170e64c7146777b05b2e93606b08db990fd5b59526f12f reviewer:critic reviewer-id:gajae-reviewer evidence:resumed-job-admission-metadata;exact-head-sdk-shard-failed;ocr-checked
…ncelled-end hang test
|
CI on 8f5efc7: session-runtime.test.ts:8636 hung again. The last persist() reached store.transact (promise transact-e8853b98) and never returned; the transact before it finished normally. The 'awaiting-lock' / 'lock-acquire' log entries are written right before store.transact, so they don't measure any lock. The stuck call is waiting on the reconciliation store's internal chain, and the test never wired the store's own instrumentation, so whatever holds the chain (most likely a transactTerminal* op) was invisible. |
probepark
left a comment
There was a problem hiding this comment.
Large PR — code review skipped; human review required.
Scope: +2092 / -293 across 34 files. OCR preview selects nine files: +846 / -79, totaling 925 changed lines. This exceeds the 800-line review limit. Even excluding native metadata leaves 923 lines. No APPROVE or REQUEST_CHANGES verdict is submitted. The PR body and previous reviews are unchanged.
Files by area:
- Async queued ownership:
packages/coding-agent/src/async/job-manager.ts. - SDK lifecycle and persistence:
packages/coding-agent/src/sdk/bus/reconciliation-store.ts,packages/coding-agent/src/sdk/host/session-runtime.ts,packages/coding-agent/src/sdk/prompt-deadline-manager.ts, andpackages/coding-agent/src/sdk/reconciliation-extensions.ts. - Task execution:
packages/coding-agent/src/task/executor.tsandpackages/coding-agent/src/task/index.ts. - Python lifecycle:
packages/coding-agent/src/tools/python.ts. - Native metadata:
packages/natives/native/diagnostic-artifact.json. - Supporting changes: queue/lifecycle/lease/ACP regression tests, read prompt/catalog and golden fixtures, and two changelog fragments.
CI: Exact-head Dev CI completed successfully. The SDK runtime shard, resume queue shard, coding-agent check, affected-path aggregate, evidence producer, and virtual integration validation pass.
https://github.com/Yeachan-Heo/gajae-code/actions/runs/38093982867
The earlier SDK CI failure is not present in this run. Passing CI does not establish that earlier ownership findings are resolved; that requires code review. No local builds or tests were run.
Conventions: Changelog fragments are present; no released changelog file is changed. The read prompt and generated tool catalog both replace the same ellipsis text. Native diagnostic metadata is changed; artifact provenance was not reviewed. No prohibited generated path or label appears in the file list.
ocr: blocking 0 / nit 0
These counts record no code findings because the size gate prevented code review. They do not certify defect-free code.
Verification: OCR preview/rule and the mandatory ocr-pr-rules.sh stamp command exited 0. The initial rule command outside the repository failed; rerunning inside the fetched repository succeeded. The run API confirms successful completion at this exact head. The pre-submit PR query confirms OPEN with the same head. The PR body's local verification predates this head and is not used as exact-head proof.
Review the production files by area, including queued-to-running endpoint ownership and SDK terminal persistence. This comment is not a merge approval.
probepark
left a comment
There was a problem hiding this comment.
Large PR — code review skipped; human review required (re-request rereq1).
The re-request follows the previous review: #6512 (review) . This response is a PR review on the requested exact head, not an issue comment.
Scope: +2092 / -293 across 34 files. Fresh OCR preview selects 9 files, +846 / -79 (925 changed lines). This exceeds the 800-line review limit. Even excluding native metadata leaves 923 lines. Re-requesting review does not remove that limit. No APPROVE or REQUEST_CHANGES verdict is submitted. The body and existing reviews remain unchanged.
Files by area:
- Async queued ownership:
packages/coding-agent/src/async/job-manager.ts. - SDK lifecycle and persistence:
packages/coding-agent/src/sdk/bus/reconciliation-store.ts,packages/coding-agent/src/sdk/host/session-runtime.ts,packages/coding-agent/src/sdk/prompt-deadline-manager.ts,packages/coding-agent/src/sdk/reconciliation-extensions.ts. - Task execution:
packages/coding-agent/src/task/executor.ts,packages/coding-agent/src/task/index.ts. - Python lifecycle:
packages/coding-agent/src/tools/python.ts. - Native metadata:
packages/natives/native/diagnostic-artifact.json. - Supporting changes: lifecycle/queue/lease/ACP regression tests, read prompt/catalog, golden fixtures, and two changelog fragments.
CI: Dev CI completed successfully on this exact head at 2026-10-10T23:46:36Z. The SDK runtime shard, resume queue shard, coding-agent check, affected-path aggregate, evidence producer, and virtual integration validation pass.
https://github.com/Yeachan-Heo/gajae-code/actions/runs/38093982867
The earlier SDK CI failure is absent in this run. Passing CI does not prove resolution of the prior endpoint-ownership finding. That requires code review. Existing peer and probepark change requests are not dismissed.
Conventions: Changelog fragments are present; no released changelog file is changed. The read prompt and generated tool catalog contain matching text changes. Native diagnostic metadata provenance is not established. No prohibited generated path or label appears in the file list.
ocr: blocking 0 / nit 0
These counts mean no code findings were recorded because the size gate prevented code review. They do not certify defect-free code.
Verification: Fresh ocr delegate preview and mandatory ocr-pr-rules.sh Yeachan-Heo/gajae-code 6512 exited 0. gh pr checks exited 0; the run API confirms completed/success at the exact head. These checks were observed on 2026-10-11 at 00:08–00:09 UTC. The PR body's older local results are not exact-head evidence. No local builds/tests or full code review were performed.
Review the production files by area, including queued-to-running endpoint ownership and terminal persistence. This comment is not a merge approval.
Fixes #6508
Root causes and fixes
agent_endhad published. It now waits for the actual terminal event; production ordering remains unchanged.agent_endto terminalize from a failure diagnostic without exact settlement evidence. The runtime now captures ownership independently of that proof seam; tests synchronize on actual observer/drain events./progressadmission test began before asynchronous session bootstrap completed. It now waits for the advertised-command update, preserving the same success, conflict, cancel, and error assertions.1a6aa6721a4bb3dd6c2ce2003c3abe465188ff72; this PR does not duplicate that change.Verification
After rebasing on
origin/devat60d0b94ca28b860e504d6068934b9541e60ec915:/progress, task artifact owner access, and release backmerge.bun --cwd=packages/coding-agent run checkpassed with four existing Biome warnings (oversized AgentSession source andnoConfusingVoidTypewarnings in unchanged Python files).—
[repo owner's gaebal-gajae (clawdbot) 🦞]