Skip to content

Latest commit

 

History

History
666 lines (540 loc) · 76 KB

File metadata and controls

666 lines (540 loc) · 76 KB
created 2026-08-02
tags
cli-jaw
codex-app
pi
opencodex
runtime-pool

Runtime Integration — codex-app / pi × opencodex

Native runtime integration의 공개 계약 요약. 풀 계약, 취소 의미론, 모델 발견 규칙, pi rpc 판정을 다룬다.

Shared event contract foundation

The shared native Cursor/Grok host admits one cached fallback start before a pre-start failure or immediate Stop emits compatibility completion. Exceptional settlement and final cleanup close only the captured still-running trace header, even if canonical recording fails. They never rewrite an already selected result, timestamp or another owner's run. Exact lease/exit-barrier cleanup remains separate; failure diagnostics are not final MESSAGE content.

src/shared/runtime-contract.ts defines native/print capabilities, distinct native-input/cancel-reprompt/queued/restart controls, and versioned presentation events. A jaw chat session and routing scope are separate from private provider session IDs. RuntimeTurnOutcome keeps authoritative finalText (null means absent; an empty string is intentional) separate from partial text.

src/agent/runtime/events.ts records a validated, redacted body through the existing trace writer before publishing agent_runtime on the agent event topic. The trace writer owns sequence allocation; sequence gaps are valid. The tuple codec in src/trace/runtime-body-codec.ts preserves numeric usage without weakening raw-trace secret masking. Known structured fragments must be sanitized before clipping by their producer. Recording failure returns null, never a fabricated event or another inference.

Codex app-server attaches projection after its existing lane/turn owner gate and legacy consumer. Host-only, foreign-thread and stale-turn notifications remain raw-only. Main multiplex ON/OFF and employee paths capture jaw identity once; lifecycle supplies the terminal rather than a tool/message completion guess. Pi RPC uses prompt-owned raw observers for start/update/end tool snapshots and accepted legacy callbacks for text/reasoning, preserving completion-only legacy tools and final echo suppression. One projection failure emits agent_runtime_gap and disables later canonical writes for that run; ordinary final delivery and salvage remain independent. No runtime/default selection changes.

Pi raw tracing omits repeated growing message snapshots, explicitly labels delta-only retention, and bounds ordinary payloads to4MiB/2048records/64KiB per record plus four small control summaries. Deliberate raw omission does not stop canonical events or legacy output; actual append failure does. Late abort acknowledgements retain the original prompt observer, while a newer prompt with no observer cannot fall back to that old observer. Pi's existing overlap rejection, probed abort and pool kill fallback remain unchanged; no new in-band steer capability is advertised. Acquisition failures always close their captured trace/result, but shared live state, status delivery, scope release, exit barrier and queue cleanup run only while the captured object still owns the scope; late failures cannot clean up a replacement.

Optional RuntimeTurnOutcome is a separate handoff for native adapters. Current Codex/Pi continue their legacy output selection when it is absent. An explicit native outcome preserves absent/empty/whitespace final values, keeps partial text for interrupted MESSAGE salvage before exit settlement, and exposes runtimeFinality/runtimeStatus, optional stopCause on compatibility/orchestrate_done terminals (including unattributed when a native ACP 130 has no kill reason), plus existing trace identity. request_settled still copies finality/status only. Public web/TUI finalization must not promote previews into a native empty answer. Messaging retains producer-owned no-response diagnostics, ACK timing and queue notices; a private native send guard rejects formatter-empty bodies before claiming delivery.

Telegram hub-member native target replies require the hub's additive bodyDelivered receipt, generated by private send observers only after successful body delivery. A later failed plaintext chunk invalidates it. The existing outbound request accepts no native guard/receipt flag; legacy callers keep their previous ok behavior. A native caller talking to an older hub without a receipt returns delivery-unconfirmed rather than claiming success or automatically sending again.

Slack's display-only subscriber can consume canonical tool events after a private RuntimeLivenessIdentity binds the admitted request to its exact run/session/scope. src/agent/runtime/liveness.ts copies identity only; the pipeline composes this notification with the collector callback, including queued runs without a collector. It does not forward raw canonical events through the legacy messaging bus or alter native finality. The final selected print result may add executionFailed:true or executionInterrupted:true to orchestrate_done; transient retries do not set them. Slack uses it for failure/progress ACK classification, independently of successful body delivery. Collector timeout/exception provenance remains separate from native finality.

Why an empty terminal cannot name its own cause

lifecycle-handler.ts builds the native outcome with lifecycleRuntimeOutcome(ctx, wasKilled || wasSteer || Boolean(ctx.stallReason)), and runtime/outcome.ts overwrites the status to 'stopped' whenever that flag is set. A watchdog timeout, a user Stop and a native steer-kill therefore arrive downstream as the same runtimeStatus, with nothing left to tell them apart.

src/orchestrator/collect.ts splits the empty-terminal fallback as far as that allows: a native runtimeStatus: 'stopped', or the legacy executionInterrupted flag when no native outcome exists, yields a stopped sentence; every other empty terminal keeps tg.noResponse. A steer still wins over both, because superseded blanks the fallback entirely — the follow-up run owns that answer (#655).

Machine runtimeStatus stays one stopped bucket. stopCause (watchdog / user_stop / steer_kill / unattributed) is the display discriminator on orchestrate_done. A raw ACP 130 without a kill reason is unattributed, not a Jaw kill. When that print-compatible exit has no native outcome, lifecycle marks the terminal with legacy executionInterrupted; pipeline then carries the cause so the collector reaches tg.stoppedUnattributed instead of tg.noResponse. The collector must not invent a cause: a missing or unknown value keeps tg.stopped. Cancellation provenance is captured with the physical or native cancellation itself and survives every runtime wrapper. Gateway replacement keeps interrupt and becomes steer_kill; explicit channel /stop uses explicit-user-stop, which keeps the same interrupt cleanup and exit-settlement behavior but becomes user_stop. The collector never infers the cause from steer_started ordering. If that event arrived first it still blanks the retired turn; if the terminal settles first, its captured steer_kill sentence is already truthful. Do not split runtimeStatus.

Captured internal reasons such as planned-restart and shutdown keep generic stopped wording. Legacy wasKilled/wasSteer flags supply a display cause only when no reason was captured; they must not relabel an internal restart as user Stop.

A runtime that DOES know why it failed now says so. ExitContext.runtimeDiagnostic carries a sentence the runtime generated itself — never relayed child output — and lifecycle-handler.ts prefers it over the stderr classification for the trace error and for the resolved diagnostic. When a failed native turn has no compatibility text, that sentence becomes the agent_done text, so the collector reports a cause instead of tg.noResponse. A stopped run stays silent and any real answer still wins. The Claude adapter fills it in settle from facade.lastError, which is where a mid-turn session failure lands; failed already had its own path and is unchanged.

Claude native cannot correlate a backgrounded task's result, so observing one ends the turn. The PreToolUse hook only refuses an EXPLICIT run_in_background: true, and Claude moves a long foreground Bash command to the background by itself, which killed turns whose command then completed normally. claude-runtime-pool.ts therefore seeds CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1 into the prepared environment when the caller left it unset; an explicit value is preserved and stays part of the pooled profile. With that switch Claude Code drops run_in_background from the Agent tool schema, so the hook treats an Agent/Task call without the flag as foreground and refuses only a flag that is present and not false (or a non-object input).

Transport selection and session identity

Independent display preference

presentation.mode is activity by default or explicitly legacy. Both fresh and upgraded documents without the field use Activity; explicit Legacy and future siblings survive merge/load/watch. This policy does not change the native-transport migration (nativeTransportMigration). API rejects invalid blocks/modes; watch ingress keeps current mode on rejected fields. A sole own presentation patch skips fallback reset, singleton session sync while preserving serialized persistence/rollback and the existing messaging dispatcher, which finds no affected transport. Mixed/empty patches retain existing behavior. Registered native ownership and live identity remain current.

Manager Display offers Activity first and Legacy as a reversible choice, with current-instance singleflight, guarded disabled edits and captured dirty acknowledgement. Classic applies a bounded generation-fenced settings refresh on settings_change, not loadSettings/runtime prompts. Failed latest reads retain the applied mode. This is preference plumbing; full Activity renderer/admission/history and TUI adoption are separate following layers, not certified by the setting alone.

Print observation and tool convergence

runtime/print-projection.ts is a counter-only accepted-content observer; print-activity.ts composes it with existing RuntimeProjection/journal bounds. Generic print and legacy Copilot ACP branches create it once. Native Codex/Pi/Cursor/Grok/Claude retain their own projections and never also take the print path. Accepted Codex phase tags, Claude deltas/snapshots, Cursor normalized segments, Grok text/thought, OpenCode steps and Kiro/AGY/Copilot accepted text are observed before destructive legacy resets. Unknown text stays unknown; stderr, housekeeping and control frames are not assistant messages. Synthetic narration/thought cards are not double-counted as tools.

The existing lifecycle's onRuntimeEnd supplies the print application-final, preserving null/empty/whitespace meaning without a native outcome. Normal and bypass error/retry paths close once. Print trace link/finalize errors are best-effort diagnostics, not rollback of an inserted assistant MESSAGE or authority for another inference/send. Reverse linkage may be incomplete after a failed trace write; body/tool blob fallback remains. This is not an atomic MESSAGE+trace-link promise.

merge-tool-log.ts keeps primary-first order, latest within each source, primary ties and terminal-over-running precedence. Identity is run+ref or run+seq, never label; unknown workers cannot use boss fallback identity. Exact-pointer parser recovery updates existing durable tool rows after RAM eviction. Only print calls RuntimeProjection.tool with allowTerminalUpdates:true; native frozen result/enrichment rules remain the default. Live snapshots read at most400 newest durable tool rows even at equal count and merge RAM fallback. The sanitizer's explicit knownOmitted option preserves a conservative known loss marker using max-overlap accounting, not addition of overlapping RAM/DB counts. Ordinary append/storage callers keep their existing additive count behavior.

Durable Activity journal

Classic restores this journal through the bounded history/discovery owners described in frontend.md. Historical scope stays as recorded even after live scope changes. Exact saved answers use the run+chat MESSAGE index, not the redacted canonical final preview or a nullable reverse trace link. A fork may read its own copied MESSAGE without gaining access to source history. Recovery never synthesizes a RuntimeEvent, answers a historical decision, or changes runtime scheduling or Slack delivery.

src/trace/activity-journal.ts commits validated runtime bodies using the existing trace sequence allocator inside one SQLite transaction. trace_runs.session_id/scope_key capture the original chat and execution scope at all provider trace starts, including internal Claude workers. The journal validates stored owner, audience and running status before append. Internal records support private child/decision lifecycle but never public SSE, discovery or replay. A copied/forked message cannot acquire source history; deleting its original chat deletes its owned traces. Additive migration backfills only trace_runs.message_id -> messages.id, never copied trace_run_id pointers, and never invents a historical scope.

Bounds: 32KiB/event, 4096 events/4MiB per run, 20,000 runtime events/32MiB globally, also bounded by configured total trace rows. A private system/runtime.control.v1 row holds append high-water, counts, close and first loss; it is not a canonical event. Integrity/storage loss stops later appends and never truncates an append-dependent prefix into a plausible answer. Control failure cannot suppress final delivery or interrupted MESSAGE salvage. Finalization closes control after an actual header update; a zero-change onlyIfRunning returns without touching completed control. Startup closes stale running records without resuming a provider.

Admission reclaims capacity from the oldest finished-run prefixes before declaring global loss. Runtime row/byte pressure expires whole prefixes and retains their ownership, high-water and explicit retention loss. Total-row pressure uses the existing raw-first retention policy. Capacity reclamation and append share one immediate transaction, so a failed insert cannot erase history. Active prefixes are never evicted for capacity; if they exhaust the budget, loss remains explicit and a run sealed by integrity/storage loss never resumes a partial journal.

Preview truncation/capacity emits the same scoped projection_degraded gap as persistence failure and records that loss in the private Activity control. Unlike integrity/storage loss, preview loss permits later bounded events and the real terminal; replay remains explicitly incomplete. A later integrity failure supersedes preview loss and seals appends. Provider execution, private I/O liveness and final answer delivery remain independent of preview capacity; no cap is raised and old lost controls are not repaired.

Discovery and replay use the trace API's explicit session query. Pages contain at most40 scanned events/256KiB and use sparse committed seq, after, fixed through, nextAfter, hasMore, incomplete and loss. A fixed high-water excludes concurrent tail appends; corrupt rows advance the scan cursor with explicit loss. Runtime rows remain immutable while legacy raw/tool rows keep their existing behavior. Retention prunes raw rows first, then whole canonical prefixes; running headers survive and loss tombstones preserve cursor meaning until eligible reclamation. Trace spill cleanup refuses symlinked roots/directories.

The existing raw drawer passes captured server session identity to summary/list/detail reads and invalidates stale/closed requests. Sparse sequence is not a row offset. This storage/API layer does not enable print projection, Activity preferences or full history UI; it cannot make historical requests actionable. Existing instance auth is retained, not a new tenant ACL. Slack final/ACK/queue behavior is unchanged.

perCli.<cli>.transport accepts print or native only for Cursor, Grok and Claude. All three default to native. Existing documents migrate absent/print once per migration id (v2 re-runs over v1 stamps); established missing-file homes and the unreadable-file stand-in also get native. Each of these uses native only where the permissions let native run. Cursor and Grok main have native ACP paths with literal auto permissions; Claude main uses the optional Agent SDK. Restrictive Cursor/Grok and unsupported worker selections are rejected with code78 before prompt-file regeneration, bucket/bootstrap/snapshot, fallback or pool work. Builtin Codex App and Pi retain their existing paths and keys. Main-adapter support and worker support are independent flags, not binary/authentication readiness.

Existing documents with no transport field are pinned to print before fresh defaults are merged, then eligible keys flip to native (nativeTransportMigration). The stamp id is native-transport-default-v2: a document carrying the v1 stamp in any state, including one whose print was chosen after v1, is migrated again once. Print chosen after the v2 stamp is kept. A v2 partial stamp retries only the engines it skipped, once permissions allow, and leaves the file untouched while nothing becomes runnable. Established missing-file homes get native transports with the v2 stamp (sessions stay on the legacy baseline); the in-memory stand-in for a corrupt/unreadable file uses native where its permissions allow, and persistence-blocked behavior is unchanged. A genuinely fresh init records factory choices explicitly. Invalid API fields are rejected; watcher input drops only the invalid transport and preserves current mode/siblings. A real watcher mode change invalidates existing ownership generations after settings commit and before publication.

Print keys are byte-for-byte unchanged. Switchable native sessions use native-v1: before the entire opaque legacy key and never update the print singleton. Spawn captures this identity once; lifecycle saves and explicit compact paths receive the captured transport/bucket. Scoped new/reset removes exact print/native keys only, including bare/default aliases for the default scope; colon-containing scopes are not a hierarchy. Instance-wide clear retains its existing Codex all-lane exception. Explicit native outcomes skip print-era automatic compact/count/high-turn-reset heuristics; native providers manage their own context. Explicit CLI switching retains its prior fresh-start semantics.

GET /api/cli-status adds runtimeSelection: {transport, nativeAdapterImplemented, nativeWorkerImplemented} only to those three engines plus builtin Codex App/Pi. These are compiled implementation flags, not authentication/binary/probe readiness. Existing status evidence and non-participating rows remain unchanged. Display preferences and runtime transport choice are separate controls; the Activity default is not enabled by this layer.

Manager runtime preference and save ownership

Manager Settings → Model defaults (public/manager/src/settings/pages/ModelProvider.tsx) exposes Runtime transport only for Cursor, Grok and Claude. Native is the default; missing transport migrates once on load, and Manager still displays a stored print honestly without creating a patch. Explicit native remains native, print is reversible, and an unknown value gets a generic error/label rather than silently selecting the first option. The unknown sentinel is UI-only and cannot be submitted. Cursor/Grok native require Auto (YOLO; permissions: "auto") and do not support native workers; Claude native supports Auto (YOLO) / Safe. The selector does not change permissions, model/effort or active CLI to satisfy those constraints, and its configured value is not a readiness check.

Three independent choices remain separate: perCli.<cli>.transport selects the next run's native/print path; presentation.mode defaults to Activity with explicit Legacy reversal; Manager preview.ts selects the HTTP embed route (origin-port, legacy-path, or unavailable none). A runtime preference edit does not select a display mode or preview route, migrate defaults, or replace the running adapter. Builtin Codex App/Pi have no print selector here.

The actual web/API settings wrapper preserves admitted-run ownership only for explicit known presentation.mode and/or eligible transport-only patches. It also leaves fallback state, singleton session untouched for these preferences; persistence, serialization, rollback and settings publication still run. Old completion saves to its captured native/print bucket; the next run reads the new choice. Mixed model/permissions/CLI/workspace or unknown/empty leaves keep the existing invalidation. Legacy presentation-only subtree side-effect skips are retained separately. External-file transport edits still invalidate ownership; the API's own saved-file echo is ignored by the existing watcher fingerprint.

components/runtime-transport-field.tsx subscribes to its own DirtyStore entry, falling back to the server original, never the row's model-draft transport. It validates CLI/value before setting exactly perCli.<cli>.transport; it sends no HTTP. The page expands valid owned perCli.*/fallbackOrder entries and uses the ordinary SettingsClient's existing PUT /i/<port>/api/settings. Save/reset share one operation owner; duplicate saves join it, ordinary inputs and already-open menu callbacks are guarded, and failed saves retain pending intent. Success acknowledges only captured entry identities, not newer or unrelated entries.

Committed client/port/store identity fences requests and completions, including A→B→A; metadata reads also have a request generation. A private snapshot read adapter tags results with their captured instance so ready A data cannot become B data; this tag never changes the wire payload or write client. Snapshot refresh overlays still-pending owned perCli values onto the model draft. These are UI currentness guards, not an authorization layer or cancellation of admitted writes.

Disabling a row retains its open Pi dialog under an inert wrapper rather than unmounting it solely for disabled state; normal snapshot loading or instance remount can still destroy the row. An already-admitted registration completion uses the separate onPiRegistered sink: current-instance provider/model intent may reconcile while inputs are blocked, but a retired instance cannot write the current draft. This does not add another HTTP mutation. Optional returned Pi profile-metadata refresh remains a separate inherited dialog response-envelope limitation; selection reconciliation does not claim to fix it. The existing Classic native-request bridge remains the embedded panel owner. This settings integration does not certify embedded browser, dev Electron or packaged-sidecar QA.

Internal Claude SDK session core

runtime/claude-sdk-session.ts owns one persistent query from the optional, exact-pinned @anthropic-ai/claude-agent-sdk@0.3.282. One reader consumes sequential parent-text turns; explicit resume is passed to a new query. The factory captures prepared options, environment and cancellation before lazy loading. The input stream has one unconsumed text slot (at most1MiB) and one active turn; this does not bound the SDK's internal buffers. Jaw sessions have no in-band steer; only a Code session (inBandSteer, set by src/code-mode/providers/claude.ts alone) accepts one follow-up per turn, described under Native Code sessions.

Turn bindings separate jaw IDs from the provider session ID. Bounded terminal dedupe, explicit user-message UUID checks and owner rechecks prevent stale identified results from completing a replacement turn. Anonymous output still relies on the SDK's single-query ordering contract. Final text remains authoritative, including empty versus absent; partial text is never promoted after error or Stop. Cleanup fences admission immediately and succeeds only after reader completion and observed owned-process closure. Native Windows launch is covered by resolver simulations, not installed-provider proof.

The internal session now maps parent streamed text, tool input/output, provider-supplied plaintext reasoning and per-turn usage through the shared RuntimeProjection. Completed block snapshots replace matching deltas; child narration and encrypted thinking never enter parent output. Tool JSON is bounded before parsing/publication. Late tool metadata may fill unknown name/input without reopening or replacing a terminal result; optional enrichment is rejected if it would erase established output under preview-budget pressure. Final publication fences reentrant input until the prior end is emitted.

The main adapter now uses this session through the existing shared runtime store, native host and lifecycle. runtime-pool-contract.ts owns type-only provider ports; the Claude adapter cannot import back into its pool owner. Prepared config/canonical cwd/environment and captured ownership govern reuse; failed physical disposal retains a fence until safe release. SDK candidates defer final publication until the host claims an immutable result and lifecycle supplies its terminal. Input remains blocked through pending/finishing state. A Stop before claim changes an unclaimed candidate; after claim the established final can survive a stopped lifecycle status, as in the common outcome contract. Error/unfinished partial is never promoted.

The internal Claude pool also accepts lifetime:'request': each acquisition gets a fresh physical query while retaining its explicit resume ID. Only forceNew clears resume. Request leases expose retireOnFinish for the native host, which awaits retirement after application settlement and before releasing the lease. Plain release also fences a request query immediately. Failed or timed-out close never authorizes reuse; the existing SDK cleanup owner retains its fence. The default pooled lifetime and its idle reuse remain unchanged.

Stop hard-closes the query, and the existing default steer policy resumes with interrupted context after MESSAGE persistence and exit-settle. Explicit followup/collect queues; jaw has no native-input hook (Code's in-band follow-up is a separate, Code-only path). No-start failure and Stop-before-acquisition use one cached fallback projection, started before compatibility completion and closed once. Exceptional trace finalization updates only a still-running header, preserving a prior lifecycle's status/timestamp/error.

Current-message partial text still resets on a new assistant message. Interruption uses a separate bounded view of the latest parent message containing a text block: a later tool-only boundary does not erase progress, but an explicit empty text block remains empty. This applies only to Stop/error and an unclaimed candidate; successful final selection and already claimed results never borrow that fallback.

An unleased ClaudeAcquireFailure.cleanup belongs to its captured main or worker control even after logical settlement removes the process-map entry. Logical answer delivery does not wait for this receipt; physical accounting ends only when it fulfills. Rejection retains the fence, and late completion never repeats lifecycle delivery. Instruction-directory cleanup remains worker-only. waitForMainProcessEnd counts main controls, including this retained cleanup, but excludes surviving workers during all three steer entrypoints. The existing waitForProcessEnd and global shutdown wait remain inclusive. Their bounded deadline returning is not evidence that physical cleanup succeeded.

Native Claude main and workers support tools, live approvals/questions, bounded image input and foreground child activity. Auto (YOLO) / Safe profiles preserve their existing meanings; deny/unknown profiles fail before prompt/directory/query work, so the output-only memory extractor still requires print. Workers use a dedicated query, real process handle and unique owned instruction directory; cancellation/completion registration outlives process-map removal until cleanup settles. Claude print remains unchanged. Foreground-only hooks do not promise an OS sandbox. SDK authentication follows the official API/cloud setup; no claude.ai login flow, credential copying or subscription entitlement is added. Qualification distinguishes actual pinned SDK/owned simulated CLI from real-provider and rendered UI evidence; these are not interchangeable.

claude-sdk-permissions.ts snapshots original input and binds callbacks to declared tool IDs. Questions return original full-question keys and comma-separated selected labels; neither auto mode nor a display item can bypass explicit ask rules. Missing/unreviewable operations deny, no future permission grant is created, and Stop/expiry cancels exactly the captured request. Images are in-memory validated PNG/JPEG/GIF/WebP, at most4,5MiB each/10MiB aggregate; the adapter never fetches arbitrary URLs/paths. Existing staged-file references remain prompt/tool access, not automatic image conversion.

Child linkage is bounded to128 children,512 tools and32 prelink frames/64KiB. Parent and child declaration paths both reconcile until no progress or32 passes. Every usable declared ID reaches the owner table even if its child finished in the same drain; live eligibility remains false for inactive children. Cross-owner/retired ID reuse fails the reader, while identical-context declarations deduplicate. Parent completion stops unfinished child display entries, never manufactures child success. A captured synchronous terminal-only capability can record old child tool endings after ownership revocation, but cannot authorize requests, text, new tools or input.

Live native decision presentation

The registry stays DB-independent with one optional change observer. Route composition installs runtime-request-notices.ts, which maps the captured registered chat to its presentation scope and emits agent_runtime_requests_changed directly to SSE. It never substitutes the active chat or calls messaging broadcast. This three-field hint contains only version/sessionId/delivery scope; canonical events and live request entries retain original execution scope.

/api/orchestrate/snapshot?session=<id> supplies activityIdentity with no-store caching and strict named-session handling. The Classic panel (also used inside Manager/Electron chat) accepts the existing same-chat live list, labels execution context and POSTs the selected row's original four IDs. Another chat or stale binding cannot be substituted. Stream health, list freshness and manual recovery are separate: initial SSE unavailability invalidates pending automatic work, manual refresh retains the outage label, only SSE-open restores live health, and an uncertain POST is never automatically replayed. Old-server WebSocket fallback is retained, not a native decision channel. Full Activity timeline/default/replay remains a separate layer.

ACP v1 transport boundary

src/agent/runtime/acp/wire.ts owns the shared single-envelope decoder. The legacy Copilot AcpClient routes a method-bearing peer request before looking up a pending client request, so equal bidirectional IDs cannot consume each other's work. Existing Copilot spawn arguments, permissions, activity timers and callbacks remain unchanged. Malformed stdout is still ignored there, without logging the raw malformed line.

src/agent/runtime/acp/connection.ts is the native transport primitive; the session/factory/main bridge owns activation and process reaping. It delivers callbacks and notifications while an outgoing prompt is pending. request() exposes dispatched (local Writable completion) separately from result (the RPC response). An id-less session/cancel write is not a remote cancellation acknowledgement: for the targeted v1 lifecycle, the original prompt response establishes completion. The newer v2 lifecycle is not mixed into this adapter. See the versioned ACP v1 transport and prompt-turn contract.

The native connection accepts string/safe-integer IDs and individual envelopes; null/unsafe-number IDs and batches fail closed. These are explicit interoperability limits, not claims that JSON-RPC forbids null IDs. Incoming UTF-8 is validated after reassembling split bytes. Payloads are capped at4MiB excluding LF/CRLF, with geometrically grown bounded carry storage (one extra byte only for a split CR delimiter). Outgoing active+queued work is capped at8MiB including LF and1024 entries; at most64 outgoing requests may await results. Each write has a30-second deadline from admission, independent of its RPC result deadline. A stalled notification/reply or an early response cannot strand dispatch indefinitely.

Malformed frames, I/O failure, timeout, EOF or child exit close once and reject all pending results and active/queued writes. Late errors are consumed; late replies/callbacks cannot reopen the connection. Diagnostics never include malformed payloads or provider error text/data. The caller must synchronously admit frames into a bounded consumer and retire/reap its own child when notified of failure. Consumer queues, real provider lifetime, permissions and remote cancel-reprompt acceptance are verified in their adapter layers, not inferred from these transport tests.

Native pending decisions

src/agent/runtime/requests.ts holds ephemeral, exact-bound decisions: runId, jaw sessionId, scope and turnId must all match, and the captured ownership predicate must remain current. It retains at most128 entries for120seconds, prunes expired/stale entries and settles once. Responses are validated synchronously before a second ownership/entry check. Invalid choices remain correctable; asynchronous validators cannot orphan the request. Cancellation data is an independent bounded, deeply frozen JSON snapshot, never a mutable caller alias or a timer closure retaining raw input.

Cancellation snapshots accept plain JSON trees only, not shared references, cycles, accessors or executable values. Copying charges JSON UTF-8 bytes before allocation grows beyond32KiB, with32-level/32768-node traversal bounds. No unrestricted clone or graph-expanding whole-object serialization precedes that check.

Admission uses the canonical request-view sanitizer and the same encode/redact/decode plus32KiB byte gate as public runtime events, including optional parent identity. The immutable stored view is cloned for GET; raw provider option identifiers and validators never enter the DTO. The response promise alone carries the mapped native decision. Restart cannot recreate executable request handles from history.

acp/permissions.ts validates core v1 permission params, including nullable titles, and chooses unattended options by protocol kind, not localized labels. Only literal auto selects allow_once (then allow_always); safe/custom arrays, including[] and['auto'], wait for a decision. acp/callbacks.ts owns at most32 callbacks per connection, while human waits remain independent of notification parsing. It maps fresh jaw handles to native option IDs only in live closures and refuses unsupported filesystem/terminal/question extensions without executing host operations.

Cancellation latches outlive a resolved registry answer. cancelRun/dispose prevent a later selected reply; cancelling an already-admitted but unflushed selected reply retires the connection because transport bytes cannot be retracted. Previously delivered bytes cannot be undone. The current-run fence rejects late grants without growing historical state or disposing the reusable dispatcher on every turn. The session adapter must dispose it on connection retirement and provide emitters/currentness bound to captured turns. Failed publication cancels the invisible request, and every mapping is disposed after its callback.

The two request routes use existing instance auth, including loopback and configured LAN bypass; they are not an OS sandbox or per-session tenant ACL. Exact IDs prevent misrouting and replay, not authority escalation between already-authorized instance users. API accepted means decision recorded, not tool completion. Provider activation, Activity controls and channel interactions remain separate layers; Slack final/ACK/queue behavior is untouched.

Grok factory boundary

runtime/acp/grok-session.ts is the internal dedicated-process factory. It reuses the existing ACP session, Windows launch resolver and owned-process cleanup. Only literal auto is admitted (--no-leader --always-approve); restrictive policies fail before spawn and keep print as the compatibility choice. Existing advertised cached-token/API-key authentication is selected without login or identity fallback. Legacy model IDs and object-valued effort choices are validated; a default alias with no explicit effort preserves the provider's current configuration.

AcpSession exposes copied setup metadata and serialized idle-only model selection. Model metadata is bounded plain JSON, acknowledgements update state before subsequent frames, and failed/aborted setup waits for owned-child reaping. The targeted provider setup requires an object response with advertised models; null-only load responses shown in the general v1 examples are an explicit interoperability limit, not invalid ACP. Grok main consumes this factory through the bridge below; workers remain unsupported and print defaults are unchanged.

The passive grok-events.ts mapper reads only aggregate _meta.usage from the original prompt response. It maps cached-read tokens without adding them to input, preserves absent versus zero, and omits malformed optional telemetry without changing the answer. Last-call counters, context size, extension payloads and cost fields are not substitutes. The existing runtime-session resultUsage hook owns event publication; main Grok activation supplies that hook in its integration layer.

Grok completion extensions do not own completion: id-less _x.ai/session/prompt_complete remains ignored, while unsupported question/plan/filesystem requests receive the common fixed protocol error. The original prompt result and callback/notification drain still gate cancellation and reuse. Captured tool updates reuse the common projector; no extra completion accumulator or native question capability is introduced.

Resident Runtime Pool (src/agent/runtime-pool.ts)

Jaw tool authorization is separate from these native pools. Qualified direct local calls to an Auto instance use full-local API authority without a per-turn secret, so ordinary tool access survives native reuse, print resume and steer. Provider approval mode remains independently captured; a Safe provider session inside an Auto instance is not an HTTP sandbox. Scoped grants remain the existing restricted path, including provider-bound RTS context. Full authority does not create missing runtime adapters or provider/account permissions.

Cursor and Grok acquisitions accept lifetime:'request' independently of native session resume. A request gets a unique physical-process key within the existing scope lane, waits for its current borrower, and retires on release. Captured options and environment precede admission callbacks. Ordinary pooled reuse is unchanged; forceNew alone clears resume.

ACP creation retains its scope entry after cancellation, deadline or replacement until the factory proves no child remains. A late returned candidate must close or produce an observed exit from that exact child before another acquisition can proceed. Explicit startup-cleanup failures without a returned child handle retain the entry: timeout, a new generation and forceNew cannot clear unknown physical ownership. There is no automatic recovery for that no-handle case. Cursor's constructor-failure path also waits for bounded physical reap before reporting a normal factory rejection. These guarantees concern owned processes, not escaped descendants or isolation from the host OS account.

Grok native main and replacement

Grok1.0.13 uses a dedicated agent --no-leader --always-approve stdio child only for literal auto. Safe/custom policies fail before spawn or prompt preparation because restrictive native enforcement is unverified. Existing advertised cached_token or xai.api_key authentication is selected without login. Legacy advertised model metadata resolves the grok-build alias and exact reasoningEfforts.value; unavailable choices fail explicitly. Usage comes only from the observed result _meta.usage, preserving absent versus zero values.

acquireGrokRuntime reuses ACP pool ownership and retirement fences in a separate engine partition; its key additionally hashes captured XAI_API_KEY, GROK_AUTH, HOME, USERPROFILE, GROK_HOME and GROK_AUTH_PATH values. Changing a supplied credential input cannot reuse an alive idle session; values never appear in the key. This is a conservative reuse fence, not proof that every variable/auth mode is supported by every installed provider version. Native session persistence uses native-v1 buckets and never print trace backfill. Workers stay disabled.

The optional common AcpReplacementTurn keeps one logical send across original cancellation, original response, notification/callback drain and idle, then replacement dispatch. Single-flight application steering returns busy/no-start for a second concurrent replacement. Fatal cancellation/dispatch/preparation/commit failures retire and never queue. The input callback runs once after local dispatch, before a fast logical final, only while the captured owner remains current; exact main identity and canonical reset generation are checked again before DB/events. Stop with valid ownership preserves an already-dispatched input fact. Returned thenables are consumed and rejected. The optional prepareReplacement callback runs after drain. Grok1.0.13 retained the tested context without copying prior input into B; this observation is not a universal provider guarantee. Anonymous packets after B begins still rely on provider ordering.

Cursor main native bridge

runtime/acp/runtime-session.ts captures one immutable application identity per send and maps ACP through the existing bounded RuntimeProjection. Active message segments are closed at real tool/thought/ID boundaries; only an eligible end_turn segment is an answer, never concatenated earlier commentary. Consecutive anonymous same-type chunks remain indistinguishable without a protocol boundary. Raw final/partial use the existing8,388,608-code-unit bound independently of3,000-character display previews and journal failure. Optional usage/observer failure does not become a new inference.

native-runtime-run.ts separates protocol result, immutable claim and application finalization. Main claims before lifecycle; onRuntimeEnd publishes the policy-selected terminal even after the main map has been removed. A private pre-broadcast attempt marker and captured outcome prevent duplicate terminal/lifecycle calls after a listener failure. Owned cancellation/cleanup/retirement/release and the captured exit barrier have independent finalization stages. Native compatibility events expose finality/status, never partialText or the full outcome object.

Server execution bindings are frozen once and forwarded through collector/pipeline/spawn. Explicit scope/chat values or captured server context survive multi-session off, including dedicated mention-watch; only automatic derivation follows the toggle. Pipeline derivation overwrites arbitrary persistedScopeId with the owned remoteKey or null, so a stale caller hint cannot select another scope. Private non-text I/O callbacks renew only the matching live collector and are removed before pooled reuse. Native live-state reads/clear compare the captured trace, so an older finalizer cannot erase a replacement. Display/history integration remains a separate layer.

Cursor cancel-reprompt

acp/replacement.ts and acp/replacement-turn.ts keep one logical send while private protocol attempts cancel and restart. The replacement waits for the original cancelled RPC response, all old callbacks/notifications, prompt finally and idle. Intermediate cancellation does not publish a logical turn-end, save an interrupted assistant MESSAGE or release the lease. A final Stop retains the existing stopped/salvage policy.

Cursor injects a captured preparation closure: original raw request plus previously accepted redirects are read-only context (the existing10-row/8000-character context budget); incomplete assistant output uses the existing4000-character tail; current operational rules remain active. Headers, separators and omission markers count against the context budget and clipping does not split surrogate pairs. Only successful input commits update the copied accepted history. Grok can reuse the neutral controller without this Cursor-specific reinjection.

The separate MainRunState.replaceTurn hook is recognized by canSteerAgent and explicit /steer. A local-dispatch callback records the original incoming text once, before a rapid logical final. Exact main object, run generation and canonical ownership are rechecked before commit; returned thenables and uncertain cancellation/dispatch/recording failures are fatal, never an automatic queued retry. Concurrent replacement admission is single-flight; busy/natural-race no-start input may queue. Stop-invalidated input is a distinct cancelled receipt and settles the existing request as cancelled without resubmission. A new-run outcome is not submitted again. /queue steer remains its separate forced interrupt/priority-run operation.

Wire limits are explicit: terminal-plus-late-content in one chunk and content in the idle gap are detectable violations; stale RPC and opaque request identities remain fenced. After B starts, anonymous same-session content carries no attempt provenance, so correct attribution relies on ACP v1 flush-before-terminal ordering. Native-input or unconditional provider memory retention is not claimed. Existing login is reused; a host using Keychain may select it with its existing credential-store environment option without changing stored credentials.

steer-input-guard.ts retains only transient pending-input cancellation tokens per scope. Native no-start handling and both slash/gateway fallback consumers hold a token until their actual enqueue decision. Scoped/aggregate Stop invalidates pending tokens even after the main run has disappeared; later new input gets a fresh set, and an old release cannot remove it. Every consumer releases in finally. This closes busy-result and return-to-consumer races without mistaking natural completion for Stop or retracting a successful dispatch.

runtime/acp/session.ts owns one child/session with initialize/auth/new-or-load, negotiated configuration, prompt response fencing, cancellation and callback/notification drain. Its matched-response observer closes new content and callback admission before later frames in the same chunk. Setup requires an actual load response; replay is not a live turn. Notifications are bounded at256/8MiB and human callbacks remain independent. Drain and cancellation each have bounded deadlines after dispatch; stopped or failed connections are never reused. Stderr is continuously discarded with only a byte count retained.

runtime/acp/cursor-session.ts launches the resolved Cursor executable with acp via the existing safe Windows launch resolver. It uses existing cursor_login authentication and no print/Copilot setup path. A startup abort has a captured-child owner; successful setup removes that acquisition-only listener. config.ts validates bounded select metadata, sets model first, reads refreshed options and applies only a supported explicit effort. An unsupported effort is an error, not an implicit no-op. In the observed Cursor build, Composer2.5 removes the effort selector; native configuration must leave effort unset for that model. An exact effort id outranks a sibling sharing its category, because Cursor advertises thinking and effort together under thought_level; a two-state false/true select is not a reasoning ladder and is excluded before that ranking. Model selection keeps strict ambiguity. A rejected model logs the advertised ids while the error stays code-only.

Print and ACP spell Cursor models differently, so config.ts accepts an opaque resolveModel hook, consulted only after an exact match fails and ignored unless its result is itself advertised. Cursor policy lives in cursor-acp-models.ts: auto becomes default, a rung is removed only when what remains is a model in Cursor's own picker vocabulary and the peeled rung agrees with the configured effort, and a version-first Claude id is reordered to vendor order. Every rule is an exact rewrite verified against the advertised set, never a similarity search; the ACP namespace advertises bare bases with effort as a separate axis, which is what makes the rewrite safe. An advertised value is never rewritten. Both axes report the same way: a rejected model or effort logs the advertised ids, identifier-shaped only and capped, while the error stays code-only.

runtime/start-failure.ts keeps the last native start failure per CLI for /api/cli-status and for the channel diagnostic. Only a run whose failed callback receives a null lease is recorded, so acquire never returning is distinguished from a teardown fault after a claimed answer; a start that reaches its lease retires the record. Only safeFailureCode hits travel, over a bounded walk of the aggregate and its causes, and an empty walk records nothing. This store is neither the cached probe row nor part of runtimeSelection. The user-facing sentence names that code and selects the model/effort wording from it rather than from a lastError substring, which a run that never acquired a lease has no facade to supply. Code Mode is out of scope: it returns its startup error to its own API caller.

acquireCursorRuntime extends the existing engine-partitioned pool. Keys include scope/canonical cwd/binary/model/effort/permission snapshot; reset generation and caller admission are checked before mutation and after creation. The total acquisition deadline aborts in-progress startup. Every borrow has a lease token: stale cancellation/retirement cannot interrupt the next borrower. Busy logical leases settle before normal replacement; release/reaper/replacement retain an entry-owned retirement fence until physical close, and a rejected close cannot admit a replacement prompt. Code/Pi policies are unchanged.

  • boss/main 실행은 메시지당 spawn 대신 상주 런타임 풀을 탄다. 키 = 엔진별 독립 스토어 + chat:${getActiveChatSession()} + cwd + 모델/effort/(pi는 profile/endpoint/apiKind/profileFp).
  • employee는 풀링하지 않는다 (매 턴 timestamped cwd + cleanup과 모순 — per-turn spawn 유지).
  • 엔트리 상태기: creating → ready(busy/dead) + 대기자 큐. 조회/마킹은 동기 임계 구역, drainWaiters가 splice→clearTimeout→resolve/reject 순서를 보장.
  • 죽은 런타임은 다음 acquire에서 재생성: codex-app은 thread/resume(복구 가능 분류 isRecoverableResumeError — 실물 에러 "no rollout found for thread id ..."), pi는 --session-id.
  • 취소는 lease의 단일 cancel(): supportsInterrupt면 interrupt(codex-app turn/interrupt {threadId, turnId}; pi는 abort 세션 계약), 아니면 kill+dead 마킹. codex-app latch 경로는 activeTurnId 부재 시 이벤트 대기(interrupt-failed/turn-completed/10s timeout) 후 실패 시 kill 폐번.
  • idle TTL 15분 리퍼. poolStats()로 진단.

codex-app exec-parity (src/agent/codex-app-client.ts, codex-app-catalog.ts)

  • thread/start와 thread/resume 모두 developerInstructions 전송 (jaw sysPrompt가 wire에 도달 — "app-server가 멍청"의 jaw 측 주원인이었음).
  • 매 turn/start에 effort 명시. model/list는 cursor 페이지네이션.
  • spawn 전 model/effort 사전검증: $CODEX_HOME/model_catalog_json의 supported_reasoning_levels 기반, 카탈로그 부재/미등재는 fail-open, 등재 모델의 미지원 effort만 fail-fast (openai/codex#31552형 행 방지).
  • 승인 자동응답: item/permissions/requestApproval에는 {permissions:{}, scope:'turn'}(빈 grant = 정상 거부 경로), 나머지는 decision/answers decline 맵.
  • pre-turn 취소 레이스는 pending-interrupt latch로 흡수 (setActiveTurnId 단일 대입점 + terminal-race 분류기, 실패는 interrupt-failed로 표면화).

모델 발견 (src/cli/opencodex-models.ts 소유)

  • codex/codex-app 모델 목록은 기존 라이브 배선(runtime-port.json → healthz → /v1/models)이 소유. resolveOpenCodexCodexModelsDetailed()가 {models, entries, source:'opencodex'|'static'}를 주고 registry가 modelSource를 노출.
  • 모델별 reasoning effort도 같은 배선으로 동기화된다. ocx는 모델마다 다른 effort 집합을 광고한다 (gpt-5.6-sol·gpt-6-sol은 ultra까지, gpt-5.6-luna·gpt-6-luna는 max까지; 2026-09-23 기준 anthropic/* routed 모델은 low..max). parseModelEntries()가 reasoning_efforts[].value / supports_reasoning_effort / reasoning_effort를 {id, efforts, defaultEffort}로 파싱하고, registry-live.ts가 codex/codex-app에 effortsByModel·defaultEffortByModel을 싣는다. efforts는 legacy 소비자용 합집합이다.
  • effort 값은 표시용이 아니라 wire 값이다(src/agent/args.ts codex 분기 → -c model_reasoning_effort="<effort>"). 그래서 UI 선택기는 합집합이 아니라 선택된 모델의 집합을 써야 하고, 빈 배열은 "이 모델은 effort 없음"을 뜻하므로 fallback 금지다. 은퇴한 ai-e만 provider 스코프 맵(effortsByModelByProvider)을 썼다. 남은 런타임은 평평한 effortsByModel이다. 해석 우선순위: 평평한 모델 → registry 목록.
  • Code 카탈로그도 같은 배선을 읽는다 (src/code-mode/providers/live-models.ts). registry-live.ts의 소비자는 설정 라우트뿐이라 Code 모드는 오랫동안 CLI_REGISTRY의 정적 목록에 갇혀 있었다. CodeProvider.describe()는 동기 함수라 ocx를 await 할 수 없으므로, 마지막으로 확보한 /v1/models 응답을 메모리에 두고 백그라운드로 갱신하는 스냅샷을 읽는다. (첫 probe가 저하되면 스냅샷은 static fallback이고, 이후 live 응답이 이를 대체한다.) read는 절대 블로킹하지 않고 이미 아는 값을 돌려준다. 규칙:
    • 저하된 probe(source:'static')는 기존 live 스냅샷을 덮지 않는다. 한 번의 실패로 routed 모델이 사라지면 안 된다.
    • 빈 목록은 저장하지 않는다. 카탈로그가 비면 validate()가 모든 모델을 거절한다.
    • 실패/저하 시도도 자체 타임스탬프를 남기고 30초 재시도 하한을 지킨다. 성공만 시계를 돌리면 stale 스냅샷 + 죽은 proxy 조합에서 카탈로그를 읽을 때마다 probe가 재시작된다.
    • 스냅샷 read는 proxy를 타는 codex-app에서만 일어난다. 필터를 나중에 하면 Claude/Cursor 카탈로그 조회가 Codex probe를 예약한다. 카탈로그는 modelSource:'live'와 effortsByModel/defaultEffortByModel을 실어 보낸다.
  • 라이브 카탈로그는 이미 도는 세션을 무효화하지 않는다. validate()는 prompt/attach/patch에서 세션의 저장된 값으로 다시 호출된다. 목록이 세션 아래에서 바뀔 수 있으므로, 생성 시점에 합법이었던 모델이 갑자기 unsupported_model이 되면 사용자가 보거나 고칠 수 없는 이유로 세션이 죽는다. 그래서 validate(input, fixed, accepted)의 accepted가 세션이 이미 쓰는 model/effort를 지목하고, 유지되는 값은 오늘의 카탈로그로 재검사하지 않는다 — 모델은 현재 목록을, effort는 현재 union과 모델별 집합을 건너뛴다. 다만 세션이 생성 시점에 저장한 capabilities는 계속 적용된다. 그것이 네이티브 런타임을 연 계약이기 때문이다. 새로 고르는 값은 여전히 현재 카탈로그로 검사한다.
  • /effort 슬래시 명령과 TUI 셀렉터도 같은 소스를 쓴다. 과거에는 ['off','low','medium','high','max']를 하드코딩해 xhigh가 아예 없고 ultra를 거부했다. 지금은 resolveEffortLevelsForCli(cli, model)이 codex 계열이면 ocx entries에서, 그 외에는 registry에서 목록을 만든다. 저장은 런타임이 실제로 읽는 perCli.<cli>.effort + activeOverrides.<cli>.effort로 간다 (top-level settings.effort는 저장돼도 아무도 읽지 않는 죽은 키다).
  • defaultModel/defaultEffort는 라이브 값으로 교체하지 않는다. buildDefaultPerCli()가 사용자 기본값을 여기서 seed하므로 ocx 라우팅 순서 변화가 사용자 설정을 조용히 바꾸면 안 된다.
  • pi 프로필 discovery: probeOpenCodexEndpointModels(endpoint)가 **healthz 핑거프린트 {status:'ok', service:'opencodex'}**를 요구한 뒤에만 <endpoint>/models 카탈로그를 사용. 불일치/미실행이면 기존 pi --offline --list-models 경로 (modelSource='pi-offline').

pi rpc 판정 (scripts/pi-rpc-probe.mts, src/agent/pi-rpc-verdict.ts)

  • 2026-08-02 실probe 판정: multi-prompt SUPPORTED + abortEffective=true (progrok/grok-composer-2.5-fast, pi 0.83.0). 판정 근거는 프로토콜 사실(두 번째 prompt의 id 상관 success 응답 + user echo 상관) — 모델이 두 번째 턴 답변을 reasoning 채널로만 내는 비결정성과 무관.
  • 판정 결과는 ~/.cli-jaw/pi/rpc-capabilities.json에 기록(schemaVersion/commandId/probedAt). spawnPersistentPiRpc는 부재/파손/schema·profile·commandId 불일치/30일 초과 시 보수적 false (cancel은 kill 경로).
  • persistent 세션은 풀 어댑터로 편입 (acquirePiRuntime); boss는 멀티턴 재사용, employee는 기존 one-shot.

Capability preparation and execution ownership

Each RPC instance now obtains one bounded asynchronous version observation, using captured command arguments, cwd and environment. Both full launch plans are validated before either execution handle is started. The actual RPC child still returns synchronously, but bootstrap/prompt writes wait for version close. The same observation selects typed settled finality and matches the existing abort receipt; both pool wrappers forward the live capability getter. Preparing prompts reserve overlap immediately, and pre-dispatch abort cancels only that reservation without sending an RPC abort. Incomplete version evidence poisons preparation; completed unknown/nonzero versions keep the existing legacy policy. The older command-availability resolver can still block for up to3s; this is not a claim that every creation operation is asynchronous.

An instance-owned cleanup controller tracks RPC and version handles. Direct execution returns a separate immutable cleanup receipt, and publishes its selected result only after that receipt. Stop uses the direct child's owned cancellation port (also used by worker/duplicate/aggregate cancellation), not a second tree timer. RPC exit starts cleanup even after version readiness. One 2s drain and1s escalation serve both handles; signals target retained handles and stop after terminal observations. Held pipes or uncertain closure retain resources, never become a successful close by detaching streams. Persistent failure/close claims its pending result before awaiting cleanup, so late output cannot replace failure with success. A previously selected typed answer remains separate from an uncertain cleanup result.

Pi workers allocate fresh temporary directories and capture canonical path and device/inode ownership. The existing release path removes only its still-owned directory with a removable physical receipt; absent allocation, missing/rejected receipt, changed directory or symlink means no deletion. Live settings do not grant deletion authority, and late close never upgrades a retained decision. These controls are not an arbitrary-process sandbox or aggregate pool/server shutdown certificate. Opaque wrappers and escaped descendants remain explicit limits; version output never enters Activity, MESSAGE or channel delivery.

Two stop ports, by definition rather than by drift

Pi joins the native runtime family through PiRuntimeSession (src/agent/runtime/pi-runtime-session.ts) with ACP-style claimTurnOutcome / finalizeTurn. turnId is the host traceRunId. cancel() is one helper, two outcomes: pooled abort-and-reuse (cancelLease); oneshot abort-then-kill. A pooled boss turn still stops through cancelTurn → requestCancel → lease.cancel() → cancelLease. A one-shot employee stops through cancelOwnedPiProcess → cancelPiExecution → oneshot session.cancel() after bindPiExecutionCancel (wrap → defineProperty → runPiTurn). runPiTurn arms cancel/watchdog first, then start(); every user prompt goes through PiRuntimeSession.send(). Employee opens with openPiRpc (no prompt write); spawnPiRpc remains a compatibility wrapper. PiLease.retire and acquire AbortSignal exist for tests this slice; spawn does not pass a signal and the stale path stays lease.release(). Do not put pooled Pi through runNativeRuntime — that runner cancels through session.cancel() rather than lease.cancel(). interrupt vs clearWorkerSlotsOnStop stays open (D7).

Pi diagnostics use captured-turn private observers. Resolved prompt stderr stays bounded in the private classifier buffer, including resolved error outcomes. Caught failures supply a redacted, bounded runtimeDiagnostic only for error outcomes; stopped turns do not acquire an error sentence. Observer failure or stale ownership cannot change the immutable outcome, and raw stderr is never added to the public outcome or Activity event shape.

Native Code sessions

src/code-mode/ owns the /api/code API. The host composes an injectable store, session manager, transcript normalizer and four direct native adapters: Codex app-server, Claude Agent SDK, Cursor ACP and Grok ACP. Each provider uses its installed CLI and existing login. Catalog availability means an executable was found; catalog reads do not start a native session or login. Sessions also persist two sidebar clocks: last_turn_completed_at (a completed or failed turn; a cancelled one does not move it) and last_visited_at (the Manager's read receipt). Neither affects runtime ownership or replay. Code Claude maps its permission picker onto the six SDK modes (ask→default, accept-edits→acceptEdits, plan, auto-review→auto, dont-ask→dontAsk, auto→bypassPermissions); init must confirm the exact mode. A permission-only change on an idle resident Claude session calls setPermissionMode on the live query and moves the approval gate instead of restarting; with no live runtime it is stored and applied on the next open. Code Claude approval cards add "Allow for this session", which returns session-scoped updatedPermissions only (never a settings file). Jaw main turns keep Auto/Safe and two-option cards. Code Claude also switches model and effort live: on an idle resident query it calls setModel when the model changed and then applyFlagSettings({effortLevel}) when the effort changed, where no effort and medium send effortLevel: null just as the open omits them. If the effort step fails after the model moved, the previous model and effort are put back, and if that rollback fails the process is retired (claude_reconfigure_inconsistent); a store write that fails after a live switch also puts the runtime back, or retires it when that fails. The binding records the new model and effort only after the runtime confirms, and a model switch drops the context usage figure. Both live switches are exclusive with prompt admission: from the busy check through the store write and any rollback the session reads busy, so a prompt, attach or second patch answers session_busy, and a runtime that is not idle refuses before any SDK call. Thinking is fixed when the query opens (on: thinking: {type: 'adaptive', display: 'summarized'}, off: {type: 'disabled'}, each with matching alwaysThinkingEnabled and showThinkingSummaries), so a thinking change, a permission change combined with a model or effort change, or a handle without live reconfiguration retires the runtime and the next turn reopens with the new options. A Claude row without a stored value reads as on. Other providers keep the restart-on-change rule. Code Claude takes one in-band follow-up per streaming turn (POST /sessions/:id/steer); jaw main and worker turns keep kill/resume steering. The follow-up is another SDKUserMessage on the live input, never with priority (the CLI default next; now would abort the turn), offered only after a top-level frame echoed a user_message_uuids array containing the primary uuid; a singular-only or absent echo keeps it not-ready. The CLI (probed on 2.1.283) folds a follow-up offered during a tool run into the running agent loop, so one result echoes both uuids, and runs one offered during a plain answer as its own CLI turn after the first result, a result per uuid. The runtime keeps the turn's accepted uuids and settles the logical turn only when results have echoed every one. An intermediate result records that segment's answer as final text and leaves the transcript, mapper, tools and notices open; it emits no usage or metadata, and the next system:init continues the same turn, so the last result carries the outcome and usage. A result with no echo while a follow-up waits fails the session claude_followup_unconfirmed rather than guessing. The turn timer re-arms once, when an intermediate result ends the first segment, so the follow-up's own CLI turn gets a full prompt window and a logical turn is bounded by two; a folded follow-up runs inside the window of the CLI turn it joined. At that boundary the mapper also retires the finished segment's message, block, snapshot, byte and tool bookkeeping, so the continuation gets the full per-turn caps (the projection keeps its items, and the last answer stays the partial text until the continuation says something); the per-turn tool-owner bound still spans the logical turn. steer() answers not-ready once 511 result ids are recorded, so terminal dedupe always holds both results. The store reserves the turn's one follow-up in code_steers (reserved, committed, rejected, unknown; a partial unique index over the three live states) before the native offer, refuses what the turn budget could not store, and commits the same-turn user_message (<turnId>:steer:<key>) only after native acceptance. A refusal before acceptance spends the key without an event; a failure after it marks the row unknown and fails the turn through the persistence path, or only answers steer_outcome_unknown when the turn had already ended. At settlement, read after cleanup, the runtime names accepted follow-ups no result consumed and their items gain phase: 'unknown'; a fold is echoed only on its result, so this means delivery not confirmed rather than undelivered. Restart recovery cannot ask the lost runtime, so every committed follow-up of a recovered turn reads the same way. A reservation still open when its turn settles or is recovered may already have been offered, so it becomes unknown (replay answers steer_outcome_unknown, never steer_key_spent); only the in-flight call that then receives the runtime's refusal records rejected. Known limitation: a follow-up's own CLI turn still passes the exact permission-mode check on its system:init, so if the CLI changed its own mode in the first segment (ExitPlanMode, a setMode session-grant suggestion) the logical turn fails claude_permission_mode_not_confirmed. Provider live-model inventory (the Cursor and Grok CLI probes) is owned by host activation — CodeHost.prime(), called once at server startup — never by a catalog read or lazy host.get().

Each backend uses code-<role>-<port>.sqlite under its own home (worker JAW_HOME, Manager dashboard home). Storage/recovery initializes lazily after HTTP binding; a failed bind cannot recover another process's live rows. Canonical Code session IDs are separate from private native resume cursors. Metadata creation is inert; prompt admission transactionally records the key, user item and starting state before native work. A repeated key returns the existing receipt, and a changed payload under that key fails. Restart marks unfinished turns as orphaned without replaying prompts. Resume failure preserves history and never silently starts a contextless replacement.

Healthy turns reuse their native handle when model/effort/policy are unchanged. Each operation captures session, turn, epoch and native handle ownership. Startup resources register before asynchronous setup and remain counted even if open rejects. Cancellation, cleanup deadlines and late exits cannot transfer ownership to a successor. closed means observed owned-resource exit/drain, not killed or logical disconnection. Claude native identity is recorded on validated root init, not deferred to the final result. Code approval policy is server-owned; available modes are provider-specific (Grok currently exposes Auto only).

A pre-preview observer retains Code content independently of the bounded Jaw Activity projection. Embedded structured data is redacted before persistence; truncation is explicit. Materialized items remain complete. Sequence-ordered code_item_update events carry append suffixes or status/phase changes, while full replacements use code_item. Production coalesces intermediate content for 50ms and flushes final/control changes. Interruption seals native callbacks, then synchronously persists already accepted pending content while the captured session/epoch/turn still owns each write. Late callbacks stay rejected; failed persistence or exhausted event capacity cannot retry content through the terminal reserve. Store settlement remains the sole turn-terminal authority. Replay and snapshots are byte-bounded; snapshots include the complete active turn or fail explicitly. Hard bounds are 4MiB/event, 8MiB/replay or snapshot, 32MiB ordinary events/turn and a separate 2MiB control/terminal reserve. Capacity errors settle as visible failed turns; real storage failures prevent success and preserve the last committed history.

Claude conversation rollback

POST /api/code/sessions/:id/rollback rolls a Claude conversation back to a settled turn by forking its native history; it is conversation only. Workspace files changed by removed turns are not reverted (no file checkpointing or rewindFiles), no prompt is sent, and the source native session is never modified or deleted.

  • Boundaries: admitTurn() mints one private prompt UUID per Claude turn (code_turns.native_prompt_uuid), and ClaudeSdkSession.send offers the prompt under it as the SDKUserMessage.uuid. It never reaches the wire. Turns admitted before the column existed stay NULL and are never targets (rollback.sinceSequence is the first user row that has one; it is stored on the session row and recomputed, from a partial index of boundary turns, only where boundaries or user rows change). A turn that settles before its prompt was handed to the runtime loses its boundary, and so does a turn restart recovery finds still starting, because a turn is marked streaming before its prompt is sent. When the native identity is replaced, every older turn loses its boundary; the first identity and the active turn, whose prompt goes to the new identity, keep theirs. The rollback commit itself is exempt.
  • Fork point: the target's prompt UUID must be in getSessionMessages(nativeCursor, {dir: cwd, includeSystemMessages: true}) and pass the human-turn-start check (so the last compaction precedes it), otherwise rollback_boundary_unavailable. The first later turn with a non-NULL UUID present in that history bounds the fork, and undispatched (NULL) later turns are skipped. A later UUID that is absent (a prompt stopped before Claude wrote it) is passed over, to the next present one or to the end of history, only when no human-turn-start entry lies between the target's prompt and that point; otherwise rollback_boundary_unavailable. When no later UUID is present at all, including when every later turn is NULL, the fork runs to the end of history under that same check. The store plan requires only the target's boundary. A follow-up reads as a human turn start, so a target turn that took one stays fail-closed there. forkSession receives upToMessageId = the entry just before that prompt, or the last entry (inclusive), and the session title when there is one.
  • Verification and remap: the fork's user/assistant bodies must deep-equal every retained source body, aligned from the end, or the fork is deleted and the answer is rollback_unavailable. Forks re-identify messages, but folded follow-ups keep their source UUID, so kept boundaries are remapped strictly by aligned position; a kept turn with no aligned message gets NULL. With the preserved-segment compaction format the fork drops the preserved messages, so any compaction disables rollback for the rest of that session.
  • Execution: claude-sdk-history-loader.ts imports the optional SDK lazily and checks each helper before use, and it is injected through the Claude provider factory. The helpers run in process, so the runtime's CLAUDE_CONFIG_DIR and CLAUDE_CODE_PROJECT_DIR_NAME must equal the server's. A transcript whose getSessionInfo size is unknown or above 32MiB is refused before it is parsed. The parse is still in-process work on the event loop; that residual is bounded by the cap.
  • Ownership: the manager checks provider, archive, cursor, the client's revision and epoch, health and busy, then reads the store plan and fences the session before its first await. Prompt (after the duplicate-receipt lookup), follow-up (after its committed-key replay), attach, patch and a second rollback answer session_busy until it ends. Inside the fence the idle resident runtime is disposed and its cleanup required, because the resident query would keep the source session. Nothing is persisted before the fork, so a crash needs no recovery. commitRollback is one transaction that compares and swaps revision, epoch and native cursor. It deletes the later items and their events, keeps their turn rows marked removed_generation (their receipts, and their follow-ups' receipts, read cancelled), leaves code_steers rows unchanged (keys stay spent, native ids are not remapped; follow-up items of removed turns leave with the other later items and follow-up rows are never targets), rewrites kept boundaries, moves the cursor to the fork, bumps revision, epoch and historyGeneration, sets the session idle with no error, and raises a replay floor. A failed commit, or a manager disposed after the fork, deletes the fork (only a fresh UUID other than the source), and dispose() waits for in-flight rollbacks.
  • Residuals: forks appear in the user's own claude --resume list, SDK-parser vs installed-CLI skew is covered only by the verification above, and an operator with transcript access can fork outside Code without it appearing in the Code index.

Retired runtime selections

Executable CLI keys exclude JWC, Claude E and AI-E. Stored jwc, claude-e and ai-e values remain readable as retired selections across boot, schema migration and settings reload; they never resolve to a different provider through a default or fallback. Execution reports retired_runtime:jwc, retired_runtime:claude-e or retired_runtime:ai-e before provider admission. Explicit selection of an available runtime restores execution; unrelated settings changes retain the retired value and existing per-CLI data. Manager and Classic show the saved retired value without offering it as a new selection.

Claude E helper discovery, agent:claude-e:* telemetry, the native/claude-e crate and the AI-E multiplexer are removed. Session buckets that still use historical claude-i / ai-e:* names are orphans; messages are not deleted.

The JWC SDK loader, engine event mapper and configuration/model-cache integration are removed. Code uses its independent native adapters. Local TUI presentation uses cli-jaw Markdown, tool rows, tree indentation, elapsed time and subagent rendering; no generated Jawcode/Bun bundle is loaded.

A new home configured with CLI_JAW_DEFAULT_CLI=jwc also keeps that retired selection and skips runtime readiness probes. An explicit supported selection in an existing settings file still wins. Unknown environment values and genuinely corrupt settings retain their existing handling.