Skip to content

feat(execd): add caller-bound command and PTY creation recovery - #1771

Open
destire-mio wants to merge 10 commits into
opensandbox-group:mainfrom
destire-mio:codex/issue-1547-idempotent-create
Open

destire-mio wants to merge 10 commits into
opensandbox-group:mainfrom
destire-mio:codex/issue-1547-idempotent-create

Conversation

@destire-mio

@destire-mio destire-mio commented Sep 8, 2026 •

Copy link
Copy Markdown

Summary

Related to #1547. If a command starts and its creation response is lost before the caller saves the execution ID, an ordinary retry starts another process. This PR adds opt-in caller-bound command and PTY creation: callers persist an operation identity and immutable request before sending; matching retries recover the original handle, and conflicting requests return 409.

  • Add /command/operations, /pty/operations, instance discovery and private operation lookup. Claim the identity in the runtime Controller before launch and return the reserved handle with 202 creating while creation is pending. Legacy command SSE and PTY creation responses retain their behavior; unsupported servers never fall back to a legacy creation request.
  • Preserve native argv, request matching, command diagnostics and authenticated status lookup. PTYs remain dormant until their first WebSocket connection and get one launch attempt per saved identity. Optional launch_attempted and launch_failed status fields distinguish dormancy, launch failure and completion after a lost WebSocket response. Reconnect, replay and takeover retain the original process.
  • Retain records, including failures, for a uniform 24-hour window. Protect active/creating records, reject new identities when capacity is full, and keep PTY lifecycle work outside the registry mutex. Configure --operation-capacity / EXECD_OPERATION_CAPACITY (default 4096); expose bounded metrics through existing OpenTelemetry.
  • Provide recovery APIs in Python async/sync, JavaScript, Go, Kotlin and C#. Instance discovery uses a 60-second monotonic cache with shared fetches and independent result snapshots. Cancellation abandons one caller's wait; mismatch/expiry invalidates future discovery without changing a saved identity or replaying creation. JavaScript/Python recovery remains an optional capability for custom command adapters.
  • Update the OpenAPI source, generated clients, recovery guide and provisional OSEP-0024. Merge upstream main through f3950db2: preserve runtime-init gating, bind operation lookup to the authenticated runtime credential, return structured recovery errors with Cache-Control: no-store, retain Go response-header capture alongside request-local headers, and move execd documentation to the new architecture directory.

The guarantee is at-most-once creation for the saved identity within one authenticated principal, resource kind and execd Controller lifetime. The window starts at the server timestamp encoded in the identity. Expired identities return 410; instance mismatch returns 409 with unknown outcome. Callers sharing an access token share its trust boundary. Command handle recovery does not replay foreground output; callers needing retained logs choose background mode before creation.

This is an experimental implementation pending maintainer design review. Memory registration and OS spawn are not transactional. Cross-controller-restart recovery, command success and exactly-once business effects are outside the guarantee.

The September 23 review is addressed in 547f082f: Go 1.20 compatibility and client-copy race, Windows startup notification ordering, C# cancellation/argv/response handling, Kotlin shared-fetch failure completion, Python exception conversion, input validation and diagnostics, browser-compatible random IDs, and OpenAPI cleanup. For the Kotlin empty-command finding, the private constructor and existing builder already reject missing/empty command and invalid argv combinations; regression tests cover those constraints without adding an unreachable adapter branch.

Testing

  • Unit tests
  • Local HTTP/process integration tests
  • Deployment e2e / native Windows verification

Validation of current merge revision 85f1a97643c73932473dc2b22e3159ffd4ba6282 on macOS arm64 (September 26, Asia/Shanghai):

  • Execd: go test -mod=readonly -race -count=1 -timeout=4m ./pkg/runtime ./pkg/web/... ./pkg/flag ./pkg/telemetry passed. This includes real child-process and HTTP recovery regressions, PTY launch/reconnect behavior, request-body bounds, capacity, runtime-credential isolation, and structured authentication/init errors.
  • Go sandbox SDK: full suite with -race -count=1, full suite with Go 1.20.14, and go vet ./... passed. The concurrent SSE/operation-lookup regression checks that operation identity headers stay on the corresponding requests.
  • Python sandbox SDK: 754 tests passed, with no failures, errors, or skips in the JUnit report. Includes asynchronous/synchronous transport error conversion, recovery, cancellation, and foreground/background stream regressions.
  • JavaScript sandbox SDK: 222 tests passed. pnpm test regenerated the API types and built the ESM/CJS/type outputs before running the suite; generated tracked sources match the committed sources.
  • Kotlin: 430 sandbox and 22 code-interpreter tests passed. Ran :sandbox:test :code-interpreter:test --rerun-tasks, including API generation and compilation, with JDK 17.
  • C#: 243 sandbox and 27 code-interpreter tests passed in Release on .NET 10, including shared-fetch owner cancellation, argv request mapping, invalid-response handling, and the foreground/background stream completion regressions introduced by the main merge.
  • Windows amd64 runtime test binary cross-compilation passed. This is compilation evidence, not native Windows execution.
  • The tracked working tree and git diff --check remained clean after validation. No additional source changes were needed for this follow-up.

The build/static checks recorded in the current merge commit were performed on September 25; the runtime/test results above were rerun on September 26 and supersede the older 63067854 test counts.

Not run in this follow-up: native Windows execution, Linux container/privileged bwrap/eBPF checks, Kubernetes/Jupyter deployment E2E, production load, instruction-level crash injection, docs build, or C# tests on additional runtime targets. Retention tests use controlled clocks rather than a 24-hour soak. Hosted CI remains separate: the 12 pull-request test workflows are awaiting approval (action_required, zero jobs); the local results above do not establish hosted CI success or maintainer approval.

Breaking Changes

  • None
  • Yes (describe impact and migration path)

Recovery is opt-in through new methods and paths. PTY launch-status fields are optional for older-server compatibility; absent fields mean unavailable launch information. Wait for PTY operation state created before attaching.

Checklist

  • Linked Issue or clearly described motivation
  • Added/updated docs
  • Added/updated tests
  • Security impact considered
  • Backward compatibility considered

@github-actions github-actions Bot added component/execd documentation Improvements or additions to documentation sdk/c# sdk/go sdk/java sdk/js sdk/python size/XXL Denotes a PR that changes 1000+ lines, ignoring generated files. labels Sep 8, 2026
@destire-mio
destire-mio marked this pull request as ready for review September 8, 2026 16:14
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-08T16:23:23.001480Z ce5ee11 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ce5ee110b1

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread sdks/sandbox/javascript/src/services/execdCommands.ts Outdated
Comment thread specs/execd-api.yaml
Comment thread components/execd/pkg/web/controller/operations.go Outdated
@kittimzhe

Copy link
Copy Markdown
Contributor

Read the OSEP and the runtime implementation end-to-end — reviewing from the agent-client perspective, since a lost create response producing a second process is exactly the #1547 failure mode autonomous clients hit in practice.

The core contract looks right: mutex-scoped claim with owner-only launch, a typed fingerprint payload (map-key sorting, absent/empty env equivalence), identity-encoded expiry instead of tombstones, and creation state cleanly separated from execution state with honest unknown-outcome semantics on instance mismatch. The failure/recovery table in the guide is the clearest statement of the guarantee I've seen in this repo.

Four thoughts, none blocking:

  1. Failed records hold a full 24h slot. A failed creation spawned no process, so the record only answers "this identity already failed" — after eviction a lookup 404s and the caller may resubmit the same ID safely (no duplicate risk, since nothing started). As written, a caller with a bug (nonexistent cwd, say) burns one of the 4096 slots per distinct operation_id for a full day, and the cap fails closed for everyone on that execd. Would a shorter retention for failed be reasonable, or is the uniform window intentional for contract simplicity?

  2. Instance lookup caching. The two-step protocol (GET /execution/instance → mint → create) is right, but the guide doesn't say whether callers should cache the instance payload. For agent loops spawning hundreds of commands, get_execution_instance() per operation doubles the RTTs. Intended pattern = one fetch per client, mint locally, refresh on operation_instance_mismatch? One sentence in the guide would prevent a lot of avoidable round-trips.

  3. Registry observability. I didn't find metrics for the registry in the runtime changes. Given the instrumentation direction in Define core metrics/traces for OpenSandbox server and systematically integrate OpenTelemetry #1766 / Define core metrics/traces for BatchSandbox controller and systematically integrate OpenTelemetry #1767, counters for claims / duplicate recoveries / conflicts / expired rejections / capacity rejections (plus occupancy and stuck-creating gauges) would make the duplicate-creation rate of real agent fleets visible — also the best evidence for whether this API gets adopted. Could be a follow-up issue so it doesn't block this PR.

  4. Minor: CreateCommandOperation launches via safego.Go while CreatePTYOperation calls createPTYSession synchronously on the request path. Intentional asymmetry (dormant session creation is cheap and lock-scoped), or worth aligning?

@kittimzhe

Copy link
Copy Markdown
Contributor

The renumber to 0024 and the follow-up in 2904c1d look great — the instance-snapshot cache (60s monotonic window, shared fetch on concurrent misses, invalidation on operation_instance_mismatch / operation_expired) is exactly what busy agent loops need, and the OTel set (execd.operation.requests by outcome, records by kind × state, capacity, and creating.oldest_age with the "not permission to retry" description) answers the visibility question nicely — including the honest treatment of stuck creations.

I read the new expireOperation path as keeping the uniform 24h window for failed records; with --operation-capacity / EXECD_OPERATION_CAPACITY now configurable the practical pressure is covered, so that works from my side.

@Pangjiping

Copy link
Copy Markdown
Collaborator

PTY operation state can never reflect a failed launch

CreatePTYOperation resolves to created as soon as the dormant session object is stored — createPTYSession can only fail on cwd resolution (components/execd/pkg/runtime/operations.go:385), so the creating window is effectively instantaneous and unrelated to process creation.

The real launch happens later, on the first WebSocket connection, with exactly one attempt (startAttempted guards in StartPTY/StartPipe, components/execd/pkg/runtime/pty_session.go:259). If that attempt fails, the WS path correctly tells the connecting client "no automatic reattempt" (writePTYStartError in pty_ws.go). However:

  • GET /execution/operation keeps returning state: created for the full 24h window;
  • GET /pty/{sessionId} reports running: false, indistinguishable from "dormant, never connected";
  • the record can never transition to failed.

So a recovering caller that follows the documented flow (wait for created, then connect) sees a healthy handle forever, while command operations do surface failed on launch errors. The asymmetry also makes the state machine misleading: for commands created means "process creation succeeded"; for PTYs it only means "session struct stored".

Options:

  1. Plumb the one-shot launch outcome back into the registry so a failed launch transitions the record to failed (the WS error path already knows the result);
  2. Or expose launch_attempted/launch_failed via GET /pty/{sessionId} so polling can tell dormant from dead;
  3. Or at minimum document in the recovery guide that PTY created covers session creation only, and a start_failed WS frame means the identity is dead and must be abandoned.

(1) keeps PTY semantics uniform with commands; (3) is the cheap fix.

@Pangjiping

Copy link
Copy Markdown
Collaborator

Caller-bound status redaction is inconsistent and drops the only diagnostic for foreground commands

GetCommandStatus blanks Content and replaces Error with a generic message for callerBound kernels (components/execd/pkg/runtime/command_status.go:73), while:

  • legacy commands on the same execd still return full Content/Error;
  • background caller-bound commands keep full output readable via GET /command/{id}/logs.

The trust boundary is identical in all three cases (any caller holding the execd token can query any handle), so this doesn't reduce exposure — it only makes the new path less diagnosable. The sharpest edge: for a foreground caller-bound command whose response was lost, the redacted Error is the only place the failure reason could surface, because foreground output files are removed when runCommand returns. Yet the message says "inspect execution output if retained" — for foreground nothing is retained. In the exact scenario this PR targets, the caller ends up with an exit code plus "command failed" and no way to learn why.

Suggestions:

  1. Drop the redaction and keep parity with legacy status and the logs endpoint; or
  2. If the intent is to keep caller-authored text (Content) out of handle APIs, keep the real Error — it is server-produced, not caller input, so hiding it has no privacy benefit;
  3. Either way, make the message distinguish background (output retained) from foreground (nothing retained) so it is not a dead end.

@Pangjiping

Copy link
Copy Markdown
Collaborator

Heads-up: this branch currently conflicts with main (mergeable: CONFLICTING).

A trial merge against latest main conflicts in:

  • components/execd/pkg/runtime/command.go
  • components/execd/pkg/runtime/command_windows.go
  • components/execd/pkg/runtime/types.go
  • components/execd/pkg/web/model/codeinterpreting.go
  • sdks/sandbox/csharp/src/OpenSandbox/Adapters/CommandsAdapter.cs

Could you merge (or rebase onto) current main and resolve these? The other touched files (spec, Python/JS/Kotlin adapters, docs) auto-merge cleanly. After resolving, a quick re-run of the execd -race tests plus the C# adapter build/tests on the merge result would be good, since the conflicts sit exactly in the caller-bound plumbing (ExecuteCodeRequest changes) and the C# adapter.

Expose PTY launch outcome through session status, retain command launch errors under recovered handles, and preserve native argv semantics and SDK request compatibility when merging upstream main.
@Pangjiping

Pangjiping commented Sep 14, 2026 •

Copy link
Copy Markdown
Collaborator

Friendly reminder: this branch currently conflicts with main (GitHub reports CONFLICTING). The conflicting files are:

  • components/execd/pkg/flag/parser.go
  • components/execd/pkg/flag/parser_test.go
  • components/execd/pkg/runtime/command.go
  • docs/components/execd.md

Could you rebase onto the latest main and resolve these? (sdks/sandbox/go/execd.go and sdks/sandbox/javascript/src/adapters/commandsAdapter.ts also changed on main but auto-merge cleanly — worth a quick double-check after the rebase.) Thanks!

@Pangjiping Pangjiping self-assigned this Sep 15, 2026

@Pangjiping Pangjiping left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Automated code review — Request Changes

Automated review (OpenCodeReview / glm-5.3) of 47ccb7af..a6ae1232 (84 files, +8394/−86): 28 findings — 1 critical, 1 high, 6 medium, 20 low.

Blocking issues:

  • [CRITICAL] Go SDK build break — sdks/sandbox/go/execution_operations.go:23,130 imports stdlib maps (Go 1.21+) while go.mod still declares go 1.20 and CI pins go-version: "1.20" (sdk-tests.yml:314). The module no longer compiles at its declared minimum toolchain; go-sdk-quality will fail.
  • [HIGH] Data race / copylocks — sdks/sandbox/go/execution_operations.go:128-130 copies ExecdClient by value (client := *e.client), duplicating an in-use streamOnce sync.Once and racing concurrent SSE callers on streamClient. Pass per-request extra headers instead of cloning the Client.

The core execd recovery state machine (claim-before-launch, 202 creating, 409/410 semantics, retention/janitor, concurrency outside the registry mutex) held up well under review. The concentrated risk is SDK client plumbing: the two Go issues above, Windows hook ordering (MEDIUM, command stays creating), C#/Kotlin single-flight cancellation/failure handling, and Python adapters skipping the SDK-wide SandboxException error conversion in exactly the network-failure scenario this feature targets.

Individual findings are posted as inline comments below. Requesting changes until the critical build break, the data race, and the six medium findings are addressed.

Comment thread sdks/sandbox/go/execution_operations.go Outdated
Comment thread sdks/sandbox/go/execution_operations.go Outdated
Comment thread components/execd/pkg/runtime/command_windows.go
Comment thread sdks/sandbox/csharp/src/OpenSandbox/Adapters/CommandsAdapter.cs Outdated
Comment thread sdks/sandbox/csharp/src/OpenSandbox/Adapters/CommandsAdapter.cs
Comment thread sdks/sandbox/python/src/opensandbox/sync/adapters/command_adapter.py Outdated
Comment thread specs/execd-api.yaml Outdated
@destire-mio

Copy link
Copy Markdown
Author

Thanks @Pangjiping for the review. I addressed the findings in 547f082f and merged upstream main in 63067854.

The fixes cover Go 1.20 compatibility and the client-copy race, Windows startup notification ordering, C# cancellation and argv support, Kotlin shared-fetch failure handling, and Python exception conversion, along with the validation and cleanup items.

For the Kotlin empty-command finding, the existing private constructor and builder reject invalid requests before they reach the adapter. I added regression tests for those constraints.

Local validation on the merge revision passed, including execd race tests, Go 1.20 tests, and the SDK suites listed in the PR description. Native Windows execution and deployment E2E remain unverified; GitHub test workflows show action_required.

Could you take another look and approve the pending workflow runs?

destire-mio and others added 2 commits September 25, 2026 17:27
Merge upstream/main at f3950db. Retain command operation recovery coverage alongside the upstream C# stream completion checks.

Validation: Go and execd builds, C# build, Kotlin source and test compilation, JavaScript typecheck and lint, and Python Ruff and Pyright passed. Runtime regression tests were not run in this revision.
Retain upstream command helpers, session error handling, set_env support,
and background process-group fixes alongside caller-bound recovery.
Resolve SDK conflicts and align compatibility fixtures with current APIs.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

component/execd documentation Improvements or additions to documentation sdk/c# sdk/go sdk/java sdk/js sdk/python size/XXL Denotes a PR that changes 1000+ lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants