Repository navigation
RFC: multi-session host (run many sessions in one process) #6247
Replies: 12 comments 1 reply
|
Review of the RFC against dev Audit vs. actual code
Global state the audit missed
ACP / SDK-broker path (how we run gjc)We run gjc through ACP from paseo on 3 Macs, about 36 concurrent slots. On that path:
Suggestions
Open question for S7: should the broker keep one host per (agentDir, cwd/worktree), or allow mixing worktrees in one host? Mixing makes the cwd/ALS work mandatory; per-worktree hosts would let some of S2 slip. |
|
Thanks — this is a much better audit than mine. Accepting all five suggestions; the RFC body is updated as Revision 1:
On the open question: my view is one host per (agentDir, worktree) first. It lets S7 ship before S2 is complete and keeps a cwd bug from crossing repos. The catch is that worktree-per-task flows (gajaeway lanes, many ACP slots) get little sharing that way, so mixing worktrees should be the second step once the ALS cwd slice lands. Happy to flip this if your fleet is mostly same-worktree. — |
Deep review of Revision 1: support the goal, not yet the implementation sequenceReviewed Revision 1, @probepark's review, and the author's follow-up against local The objective should remain lower total memory for equivalent completed work, while preserving existing runtime, harness, SDK, daemon and Broker contracts. A multi-session host is a candidate mechanism, not the objective itself. Faster warm startup is useful secondary evidence, but should not silently replace the memory goal. Moving measurement to S0, adding the Worker comparator, expanding the inventory, prohibiting shared-PID session reaping, and adding session diagnostics are substantial improvements. I would not repeat those as missing features. Remaining verdict: ITERATE before implementation admission. In particular, the follow-up's proposal to ship S7 before S2 is complete is not safe merely because sessions share a worktree. 1. Remaining blockers and qualificationsA. Same-worktree grouping is an optimization boundary, not session isolation. Two sessions in the same Full mixed-worktree support may be deferred. Every path reachable by the first hosted workload must nevertheless have isolated ownership, or an explicit, tested admission rejection before execution. A narrower first rollout must not alter the flag-off runtime's supported behavior. Do not silently break moves or launch a hosted session that can later invoke an unsafe path. B. ALS is a context carrier, not an automatic isolation proof. Using the existing ALS conventions is sensible; replacing all call sites is not necessary. But each boundary needs a proven owner: SDK requests, socket callbacks, event handlers, scheduled jobs, queued work, child/Worker messages, startup, teardown and late callbacks. Registering a callback while a session context is active is not by itself proof that later event dispatch restores that session. Concrete examples:
Tests must include A/B callback interleaving, identical job IDs in unrelated sessions, disposal of A while B continues, and a delayed callback from old A after its identity is reused. Missing context in an explicitly session-bound operation must fail closed, not choose another live session. C. S3's "maps → per-owner instance" is not justified for every map.
Classify by resource/authority, not merely module scope. Keep a canonical-resource coordinator shared where required; session leases identify callers and retirement. Test same-resource contention across sessions and processes, different-resource independence, path aliases, late release and successor preservation. A concurrent-release smoke test is insufficient. Similarly, D. Session close and host termination require separate authority and separate proof. The new no-shared-PID-reaping rule is necessary, but S7 still bundles too much. Existing lifecycle close/startup rollback uses incarnation checks, readiness/effect markers, exact unregister/cleanup evidence and honest Distinguish:
A close timeout for A must never escalate to TERM/KILL of the host as though the PID belonged only to A. If A cannot be proven quiescent, retain uncertainty and authority; do not admit a replacement on a promise that cleanup probably finished. Any eventual host-level termination is a different, explicitly authorized action affecting every hosted session. E. An in-host event-loop watchdog cannot rescue a hard synchronous hang. For the shared-main-event-loop variant, a timer cannot run while the loop is blocked. ALS does not change this. Use an observer outside that loop/process to detect lost progress. Native aborts, runtime crashes and OOM remain host-wide failures; do not promise session-level containment for them. Start with hosted capacity 1, then 2 in tests/dogfood, and only promote to 4 after memory, fairness, latency and failure evidence. Use host-wide resource budgets and bounded admission, not just a session-count limit. Soft lag can stop new admission and drain at safe boundaries; a stuck host cannot transparently live-migrate its in-memory work. F. Re-home is not permission to replay an uncertain turn.
Fault injection must cover before acceptance, after acceptance/before persistence, after side effect/before terminal publication, and during close/unregister. A host crash should affect its own sessions only, but the error must honestly expose the increased correlated failure radius. G. The Worker alternative is promising, but is a different design branch. Bun documents separate JS instances/event loops and says Worker Bun also documents Worker termination as experimental, and H. Harness, daemon and TUI are acceptance surfaces, not incidental call sites. The harness currently creates a PTY-backed child and owns bounded TERM/KILL teardown ( The SDK lifecycle ownership contract already says Broker owns session lifecycle and provider supervisors own transport ( Process signals have no session ID. OS TERM/INT is host-level; targeted session abort/close should use the existing authenticated control surface. TUI Ctrl-C, raw terminal ownership, resize/restore and stdout must keep their current semantics. LSP module handlers currently call 2. Recommended sequential micro-slicesThese are dependency checkpoints, not authorization to implement. Each numbered item should be a small independently green PR; split further rather than satisfy a file-count limit by hiding a semantic change in a proxy. Apply only the changes required by the execution model chosen at step 4: Worker-local isolation can satisfy some main-loop context/refactor checkpoints without a rewrite, but that exemption must come from the inventory and tests, not an assumption that Workers isolate everything. Phase A — prove the opportunity before changing the foundation
Phase B — ownership inventory and narrow context contracts
Phase C — retire session resources without touching neighbors
Phase D — change lifecycle topology in small steps
Phase E — integrate every current product surface, then expand
Independent investigations can happen in parallel. The load-bearing implementation checkpoints should merge sequentially on a working product; no merge depends on an unfinished later slice to restore functionality. 3. Acceptance and stop gates
Finally, reconcile the actual Slices, Regression gate, and Kill criteria sections with the Revision 1 header/reply. They still describe the old S1–S8 order, per-owner lock rewrite, six-handler consolidation and measurement only before S7. One authoritative plan is needed before an implementation PR; the accepted corrections should not exist only in a summary banner and comments. Bottom line: measurement-first and incremental migration are the right direction. Prefer the least invasive model that demonstrably meets the memory target. Do not treat ALS, same-worktree grouping, a watchdog or a passing single-session smoke test as permission to skip session ownership and lifecycle proof. |
Follow-up: dev integration strategy is part of the safety designOne additional concern from the owner: this work may run long enough that keeping a large migration branch synchronized with an actively changing I agree. Define the big picture once, but deliver it as small, independently complete behavior/refactor PRs merged progressively into Recommended delivery model
What every slice PR should declare
Handling changes on devBefore opening a slice, and again before merge, inspect upstream changes to its contract and callers. When Any changed PR head needs fresh verification and the required approving write-access maintainer review on that exact head. Do not treat an earlier approval as surviving a rebase or substantive fix. The roadmap should capture a prerequisite that moved, but it should not freeze unrelated development while the migration runs. Prefer semantic boundaries over an arbitrary file count: for example, one local-URL ownership correction, one Settings hook/lifetime correction, one child-env launch chain, one session-close proof change, or one harness launch/stop integration. These may be much smaller than an original S1/S2/S5/S7. Splitting a large change into several PRs must not hide its cumulative blast radius or bypass re-review. Why this matters specifically hereThe risky files are central and actively consumed: Settings, session construction/disposal, endpoint authority, Broker close/recovery and harness transports. A long-lived replacement branch duplicates that moving foundation. Late rebases then combine unrelated production fixes with an architecture migration, making it difficult to distinguish migration regressions from ordinary upstream changes. Progressive merge avoids that divergence, but is not permission to merge incomplete behavior. At every checkpoint:
Proposed amendment to the RFC: make progressive integration into |
Revision 2 review: Phase A is supportable; this is not implementation approvalRe-reviewed the authoritative Revision 2 body (updated Updated verdict: OKAY for the admitted Phase A research only. The earlier blanket ITERATE verdict should not be carried forward as though the author had ignored the review. Revision 2 resolves the major planning objections:
This is now a credible incremental direction. I do not see a reason to block baseline characterization and measurement. The following are concrete Phase A clarifications and later admission requirements—not demands to implement B–E during research. 1. Correct the new per-module/Rust attribution claimPhase A.2 now says a “per-module memory breakdown” ranks Rust-port candidates. Treat that as exploratory evidence, not a causal or additive accounting of process memory. A module can create objects retained by a different owner; importing it can initialize shared runtime/native services; allocation location is not necessarily the lifetime owner. Allocator residency, native buffers and Worker heaps are separate from module source/bundle size. A module-size or import-cost table does not establish that rewriting that module in Rust saves its alleged RSS share. Require attribution labels and controlled counterfactuals: what was enabled/disabled, which retained owners changed, whether work/output stayed equivalent, and where bytes moved. Reuse the existing policy in Suggested edit: “Capture owner/allocation/import evidence where measurable to identify hypotheses; rank targeted optimization candidates only after representative counterfactual evidence. Per-module estimates are not additive RSS or native-port approval.” 2. Make the Phase A comparator executable and honestUse the same completed work, capabilities, transcript sizes, tool/subagent concurrency, cancellation and cleanup expectations for standalone versus Worker candidates. Include the host, Workers and all descendants without double-counting. A minimal Worker running a reduced synthetic workload compared with a fully initialized standalone CLI is not an architecture win. The repo's A bounded prototype may legitimately omit capabilities, but then label its result opportunity evidence, not a validated model preserving the complete product. Warm Worker creation/open is also not semantic SDK readiness. Worker isolation exemptions must be tied to actual tests on the pinned runtime, not just a separate module graph. 3. Resolve the research/prototype boundary without opening a migration branchPhase A permits measuring a Worker candidate while Status prohibits implementation PRs until numbers/model selection. Clarify that the permitted work is a disposable bounded experiment or measurement fixture, not production lifecycle/Settings rewrites disguised as prerequisites for a benchmark. Specify resource/time limits and externally supervised cleanup. Do not let experimental startup allocate real user sessions, replace endpoint authority or retain orphan tools after the run. If the full production path cannot be exercised without unsafe foundation changes, report that limitation and keep the model decision provisional; do not quietly perform B–D inside A. Keep the four Phase A activities as separate evidence deliverables where useful: current behavior inventory, standalone baseline, bounded comparator, then the written decision. This preserves the “no long-lived replacement branch” requirement even during investigation. 4. Publish the numerical decision contract before candidate resultsThe plan correctly requires advance thresholds but does not yet provide the numbers. Phase A admission should produce the decision protocol before the candidate runs are inspected:
Do not invent a percentage after seeing a promising result. Use standalone calibration to characterize variance, then freeze candidate acceptance. A quicker startup with no meaningful memory benefit does not satisfy the primary objective. 5. Before affected implementation slices, instantiate the matrix—not just its headingsThe acceptance gates are much stronger now. Turn them into concrete existing/new cases when each slice is admitted, covering standalone CLI/TUI, ACP, native/loopback SDK, Broker lifecycle/restart, provider-daemon routing and harness PTY/stop ownership. Bind results to exact base/head and compare test identities/statuses. For checkpoint 19, For checkpoints 15 and 19–23, specify what proves the session instance fully retired, including late resource authority, independently of host death. Today Production routing to capacity 2 must wait for the applicable identity/cleanup, crash-reconciliation and surface gates. A two-session test harness can precede those later integrations; a live product workload cannot bypass them because “capacity 2 already passed.” Final admission boundaryProceed with characterization, measurement protocol and the bounded Phase A experiment. Do not infer approval for foundation mutation, production pooling or rollout. Publish the numbers, limitations and model decision here, then admit only the next necessary contract-sized slice from current The correct destination is not “finish all 26 changes.” It is “meet the memory target with the least invasive proven changes, while every merged head remains a working product.” If the measurements justify stopping after tuning or a narrower optimization, that is success—not an unfinished migration. |
Additional critic pass: live dev upgrades and opt-out after Broker restartA focused follow-up with the critic found two new lifecycle admission details worth adding. These do not reverse the Phase A-only research verdict, and they do not repeat the earlier measurement/ALS/cleanup comments. Source review only; no tests or benchmarks were executed. 1. Warm-host reuse needs loaded-code generation, not just a live PID/version/pathProgressive merging into Existing mechanisms already protect important boundaries:
However, package generation is the package version ( This is not a claim that the existing standalone Broker is broken. It is a requirement for the proposed new warm-host allocation path: a fresh caller must not attach a new-generation session to an old host merely because its PID, version string and executable path still look valid. Amend checkpoints 19–22 to specify:
The concrete tests are: same package version with changed source; old Broker/old host/new caller; Broker restart with both old and new hosts present; same-path executable replacement; and a late old-generation unregister after a new allocation. An on-disk Git HEAD is not proof of which code a running source process actually loaded. Also measure the upgrade overlap: old draining hosts plus newly started hosts can temporarily duplicate precisely the runtime cost pooling aims to remove. Include a long-running or cleanup-uncertain old allocation; global budgets must prevent uncontrolled new cohorts without falsely claiming the old one has retired. 2. Opt-out controls new admission; it must not erase existing allocation ownershipThe RFC correctly says disabling pooling stops new admission while existing sessions drain. Make this durable across restart, not just an in-memory branch:
If the restarted Broker chooses the default standalone executor solely from the current flag, it could lose the ability to reconcile the hosted allocation or apply the wrong process-level teardown to it. A failed/ambiguous hosted admission must also never be retried as a new standalone spawn just because the flag changed. Existing managed-spawn recovery is a useful precedent, not evidence that the new pool topology is already implemented: Amend rollback/checkpoint 23 to require:
Test both flag transitions around startup acceptance, close and uncertain cleanup, with Broker restart between the steps. Assert no duplicate work, no shared-PID signal, no successor removal and no lost retirement authority. These are narrow additions to the progressive-dev strategy: each new merged generation must know which hosts it may allocate into and which pre-existing allocations it still owns the obligation to retire. They are later implementation-admission contracts; Phase A measurements remain the only currently admitted work. |
Phase A measurement boundary: ownership graph, not just a parent-process treeA second bounded critic pass found one concrete measurement issue not covered by the previous general RSS cautions. Real GJC processes can deliberately leave the benchmark driver's process tree. This matters before any Phase A numbers are accepted. Source evidence (inspected, not executed):
This is not an existing fixture-sampler bug. Its declared synthetic scope is appropriate. But using the same driver-rooted PPID walk unchanged for the proposed real concurrent sessions could omit the detached/reparented Broker, session hosts and their tools. That would make “total memory” incomplete. Conversely, summing one process tree per session could count a shared host or Broker repeatedly. Bun Workers within one host contribute to that host process's RSS; they are not extra independent process-RSS totals. Add this acceptance condition to Phase A:
The collector may use existing discovery/index evidence without changing production lifecycle authority. This is a measurement-only prerequisite, not justification for new pool machinery during Phase A. The critic found no other new actionable item in this bounded pass. Verdict remains Phase A research only; a model decision requires complete, explicitly scoped ownership accounting in addition to equivalent work and the previously requested numerical gates. |
|
This RFC is a tough read given the sheer volume of global state involved. The reference comment's point about I’d strongly agree with the suggestion to use For the "no scaffolding" rule, I’d be careful about the Bun Worker variant. While it isolates globals for free, the IPC overhead might negate the memory savings if sessions are chatty. Maybe start with a small PoC on a subset of sessions before committing to that architecture. On a practical note, when I’ve been dealing with similar complex state transitions in distributed systems, using AI Pro from aipro.dev to help trace dependency graphs and simulate edge cases has actually saved me a lot of manual grep time. It’s not a replacement for the deep review here, but it helps spot those hidden global mutations faster. But definitely, the 15-file kill criterion is smart—force small, verifiable slices. One question: for the ACP fleet, do you have visibility into the actual CPU contention when multiple sessions share an event loop? A sync hang in one session freezing others is a real risk. Maybe the watchdog for event-loop lag should be a prerequisite for S7, not an afterthought. |
|
This discussion is about RFC: multi-session host (run many sessions in one process). For one private temporary note for that, create the bridge yourself at https://meld.mergeinc.workers.dev. Say what it is for and what it is not for, and send the URL privately. It stays open 36 hours until the first reply. Each later reply resets 24 hours. The conversation stays on that same link. It is free, and there is no billing. It is not session identity, and it will not resume the old session. |
Phase A pre-registration (posted before any gated run)The harness is in PR #6318 (bench-only, with no changes under
Preflight and the Worker arm refuse to run unless the receipt matches this digest. Any change to the contract text changes the digest and invalidates this receipt. Arms and accounting
GatesAll gates are evaluated at N=5, using the median of at least 5 valid repetitions per arm. Fewer than 5 valid repetitions gives
Claim rule: an unqualified all-gates Preflight (stop rule)Two Workers run in one host, and session 0 is disposed while session 1 has an in-flight model turn. The turn interval must span session 0's Known runtime hazard (disclosed before running)About 2 in 20 Worker smoke launches on Bun 1.4.0+34cbb9a40 crash with Results, raw-artifact provenance (git SHA, Bun version, macOS build, CPU), and the verdict will be posted in a follow-up comment. Appended 2026-10-04T14:33Z, after the runs, for completeness; the gates above are unchanged. The contract JSON below is byte-identical to preregistration.json{
"schemaVersion": 1,
"experiment": {
"rfc": "#6247",
"phase": "A",
"scope": "bench-only isolation-model memory experiment",
"arms": ["standalone", "worker"],
"comparisonConcurrency": 5,
"brokerFreeArms": true
},
"thresholds": {
"minimumValidRepsPerArm": 5,
"memory": {
"workerToStandaloneMeanMaximum": 0.5,
"sampledPeakRegressionAllowed": false,
"concurrency": 5
},
"turnLatency": {
"maximumRelativeIncrease": 0.1,
"percentile": 0.95
},
"throughput": {
"minimumRelativeChange": -0.1
},
"eventLoopLag": {
"maximumP95Ms": 50,
"percentile": 0.95,
"aggregation": "maximum of each thread's p95"
},
"coldReadiness": {
"maximumRelativeIncrease": 0
},
"teardown": {
"workerHostFootprintMultiplierOfBrokerBaseline": 1.15,
"sampleDelayAfterFifthCloseMs": 10000
},
"churn": {
"cycles": 20,
"maximumCycle20GrowthOverCycle1": 0.1
},
"orphans": {
"maximumOwned": 0,
"maximumUnresolvedOwnership": 0
},
"fidelity": {
"required": "equal"
}
},
"definitions": {
"steadyStateWindow": "Every scheduled 1Hz sample tick from 2000ms after the last session is ready until the first session starts disposing; this spans the active workload phases and the scripted idle between them, because mock-provider turns complete in milliseconds and a 1Hz sampler cannot isolate them. The window denominator is all scheduled ticks.",
"sampleRateHz": 1,
"sampledPeak": "Maximum cohort total at a scheduled sample instant within the steady-state window.",
"p95TurnLatency": "Per repetition, nearest-rank p95 of endedAt - startedAt over all RunnerEvent turn events from all sessions; the gate compares medians of per-repetition p95 values.",
"throughput": "Per-repetition count of completed ok turns whose timing is within the active window, divided by active-window seconds; the active window starts at the first active:* phase and ends at idle or disposing after the last active phase.",
"eventLoopLag": "Per repetition, compute nearest-rank p95 of 100ms setInterval drift samples separately for every runner thread/session, then use the maximum thread p95; the gate requires both arm medians to be at most 50ms.",
"coldReadiness": "Per-repetition median of session coldReadyMs; the gate compares arm medians. Worker first-session readiness includes host process spawn.",
"warmWorkerAdmission": "Reported separately and is not a gate.",
"teardown": "Worker host process footprint from the complete sample taken 10000ms after the fifth close, compared with 1.15 times B. Standalone uses the same post-close sample time for residue evidence.",
"churn": "One resident host completes 20 cycles of create×5, full workload with idle 0, and close×5; each footprint is sampled 10000ms after the fifth close. Worker host footprint at cycle 20 must be no more than 1.10 times cycle 1. Standalone uses the protocol for orphan/residue comparison only.",
"armTotal": "Sum of physFootprint over owned incarnations present at the tick, including the driver incarnation exactly once. The driver's own footprint remains a separate diagnostic; no Broker is an arm member.",
"gatedTotal": "armTotal + B, adding the isolated Broker-only baseline exactly once to each arm.",
"memoryComparison": "At N=5, median Worker gated-window mean must be at most 50% of median standalone gated-window mean, and median Worker sampled peak must not exceed median standalone sampled peak.",
"memoryFormula": "For each arm, gatedTotal(t) = armTotal(t) + B, where armTotal already includes the driver exactly once; pass the 50% mean threshold only when W + B ≤ 0.5 × (S + B), with no sampled-peak regression.",
"contractDigest": "SHA-256 hex of canonical JSON for this file; the digest is external and is not embedded in the immutable contract.",
"postingReceipt": "Store contractDigest, discussionCommentUrl, and postedAt separately in preregistration-receipt.json before preflight or Worker admission.",
"brokerBaselineB": "Median across at least five complete repetitions of each repetition's mean 1Hz footprint samples over 30 seconds (shorter diagnostic baselines are never eligible), after trampoline exit and idle readiness; acquired from the isolated ensureBroker discovery pid and incarnation, then the exact owned Broker is stopped and verified dead.",
"statistics": "Each gate uses the median of per-repetition measurements from valid N=5 repetitions; an arm with fewer than five valid repetitions makes that gate insufficient-evidence.",
"percentileEstimator": "Nearest rank: sorted[ceil(p × count) - 1].",
"repEligibility": "A rep is valid only when every scheduled gated-window tick has exactly one complete sample with a resolved total; its 10000ms post-close teardown sample is complete; visibility self-check passed; every session has all REQUIRED_WORKLOAD_EVENTS; and invalidReason is absent. Zero-tolerance completeness applies: missing ticks are never zero-filled or omitted.",
"sampleCompleteness": "Complete only if every owned member has a resolved footprint read, no owned read is unresolved or raced, no unresolved-ownership process exists, and the driver read is resolved. A non-driver member the OS confirms exited between the ownership scan and its read holds no memory and contributes nothing without making the sample incomplete. Incomplete samples remain in raw output but their partial totals are never used.",
"churnEligibility": "All 20 Worker cycles must complete with no session failure, and every cycle post-close sample must be complete; otherwise churn is insufficient evidence.",
"orphanEligibility": "Both owned-process and unresolved-ownership orphan counts must be zero; a surviving process whose identity cannot be read or whose ownership cannot be proven counts as unresolved. Every orphan receipt must be complete (process enumeration and ownership scan succeeded); an incomplete receipt is insufficient evidence.",
"fidelity": "Normalized runner evidence of every Worker session must equal the same (N, repetition, session index) standalone session; a missing repetition or session on either side is a difference. Real-host characterization is separate and only caps the claim label (capabilityLabelRule)."
},
"environment": {
"platform": "darwin-arm64",
"bunVersion": "record Bun.version in run and preflight provenance; Worker admission is bound to the current Bun version",
"osVersion": "record the macOS version in run environment pins",
"bunConfiguration": "default configuration",
"smol": false,
"deviationFromRfcA3": "The experiment uses default Bun configuration and does not use --smol; this is the recorded deviation from RFC A.3.",
"configurationBinding": "The bootstrap digest binds the effective runner/bootstrap configuration and is checked against preflight provenance."
},
"bindings": {
"workloadDigest": {
"runtimeValue": "computed by workloadDigest() outside this contract",
"checkedAgainst": "current run inputs and passing preflight provenance",
"digestValueStoredHere": false
},
"bootstrapDigest": {
"runtimeValue": "computed by bootstrapDigest() outside this contract",
"checkedAgainst": "current run inputs and passing preflight provenance",
"digestValueStoredHere": false
}
},
"cohort": {
"ownership": "Unique GJC_BENCH_COHORT environment marker is primary evidence; retained Darwin ancestry is a cross-check and detects marker-scrubbing descendants.",
"visibility": "Only a complete KERN_PROCARGS2 environment read without the run marker proves marker-free status. Unproven visibility without independent exclusion is unresolved ownership and incomplete evidence.",
"limitations": "A descendant that scrubs its environment and is born from an unobserved intermediate that has already exited cannot be detected; the fixed workload forbids env -i-style commands and bootstrap rejects scripts containing them.",
"cleanup": "Supervision records pid and incarnation for each root and terminates by exact identity on timeout or abort."
},
"broker": {
"membership": "Broker is not a member of either arm cohort.",
"charge": "Add B exactly once to each armTotal; do not add the driver separately because it is already included in armTotal."
},
"workerPolicyException": "This is a bench-only Worker exception. No production Worker-policy or compile-entrypoint registration is changed; the Worker entrypoint remains bench-local, and the experiment makes no compiled-binary Worker-entrypoint claim.",
"capabilityLabelRule": {
"equal": "full-capability",
"differsOrUnavailable": "isolation-model opportunity evidence",
"claimRule": "Only an equal characterization permits a full-capability or all-gates-pass claim; differs or unavailable caps the report label at isolation-model opportunity evidence."
}
} |
|
Closing per maintainer decision: not actionable at this time. Can be reopened if the need resurfaces. — |
Phase A results: stop at Phase A (memory)This was measured against the pre-registration above (contract digest Gates (N=5; 5/5 valid repetitions in each arm)
The capability label is isolation-model opportunity evidence: characterization is What the numbers say
Disclosures
ConclusionPer the pre-registered stop rule, Phase A ends here, and Phases B–E are not admitted. With Bun Workers, the multi-session host's isolation model does not deliver the memory opportunity the RFC hypothesised (≥ 50% saving at N=5). Worker churn is also not stable on Bun 1.4.0. Reopening this would need a different sharing model, not a refactor of the Worker approach:
Appended 2026-10-04T14:33Z:
|
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Objective
Lower total memory for equivalent completed work, while keeping the existing CLI/runtime, SDK, Broker, provider-daemon and harness contracts intact. A multi-session host is one candidate mechanism, not the objective. Warm-host spawn-to-ready latency (ACP fleet today: p50 ≈ 38s, p90 ≈ 88s, n=60) is tracked as secondary evidence and never replaces the memory goal.
Delivery model (applies to every phase)
dev. No long-lived migration branch. Each PR is a complete vertical change cut from currentdevand merged on its own; nothing merged needs a later PR to be correct.Every slice PR declares: purpose and boundary (one contract, non-goals, merged prerequisites); exact base/head SHAs and affected entry points; merge-time acceptance; standalone value at this head; operational reversal; and memory accountability (measurement, enabling work, or measured reduction).
Phases
These are dependency checkpoints, not implementation authorization. Each phase is admitted here before work starts.
Phase A — prove the opportunity (no foundation changes)
smolincluded fairly. Agree on the memory-win and latency/throughput thresholds before looking at candidate numbers.Phase B — ownership inventory and narrow context contracts
Settings.loadForScope), then separately prove durable writes, listeners, default-profile updates and borrowed/owned close lifetimes.sdk/host/control/runtime-gate.ts), startup metadata and diagnostics; stale or wrong-owner access is rejected.move_session,/move,session.cwd.move) preserved in their own slice. Mixed-worktree admission may be deferred; safety for reachable moves may not.Phase C — retire session resources without touching neighbors
local:///agent:///artifact://resolution; no cross-session visibility or cancel, and no last-created-session fallback.inFlightget an explicit host-wide vs per-session decision.Phase D — change lifecycle topology in small steps
Phase E — integrate surfaces, then expand
Acceptance gates
docs/perf-profiling-corpus.md.Stop and rollback
dev.Status
—
[repo owner's gaebal-gajae (clawdbot) 🦞]
All reactions