Skip to content

refactor(server): make snapshot create path synchronous for k8s/fsb runtimes - #2098

Open
Pangjiping wants to merge 1 commit into
opensandbox-group:mainfrom
Pangjiping:refactor/snapshot-synchronous-create-k8s-fsb
Open

Pangjiping wants to merge 1 commit into
opensandbox-group:mainfrom
Pangjiping:refactor/snapshot-synchronous-create-k8s-fsb

Conversation

@Pangjiping

Copy link
Copy Markdown
Collaborator

Summary

  • Root fix for the early-read Failed race (refactor(server): make snapshot create path synchronous for k8s/fsb runtimes (root fix for early-read Failed) #2092, follow-up to fix(server): don't fail a snapshot that is read before its worker runs #2077): for k8s/fsb, POST /sandboxes/{id}/snapshots now creates the runtime object (SandboxSnapshot CR / fsb submit) before persisting the Creating row and returns without touching the worker pool. Both submits are fast and idempotent (CR name / fsb request_id derive from the snapshot id), so once a row is readable the runtime object always exists — read-time sync and recovery can no longer map a not-yet-created object to Failed, including on other HA replicas.
  • Recovery resubmits stuck k8s/fsb rows inline instead of queueing workers, keeping the worker pool and the _worker_in_flight guard docker-only (the read-path guard is removed; it was unreachable for docker rows, which have no change stream).
  • A failure to persist the row after runtime creation triggers best-effort deletion of the just-created runtime object; a failed submit returns 500 SNAPSHOT::RUNTIME_CREATE_FAILED with no row persisted.
  • Docker keeps the existing async worker pool unchanged (container.commit() is genuinely slow).
  • Docs: updated the snapshot create ordering in docs/architecture/fast-sandbox/checkpoints.md.

Closes #2092

Testing

  • Not run (explain why)

  • Unit tests

  • Integration tests

  • e2e / manual verification

  • uv run ruff check clean; uv run pytest full suite: 2173 passed (with OPENSANDBOX_TEST_POSTGRESQL_DSN, includes the 3 HA PostgreSQL integration tests; pyright has no new findings vs main).

  • New/updated tests cover the acceptance criteria from refactor(server): make snapshot create path synchronous for k8s/fsb runtimes (root fix for early-read Failed) #2092:

    • sync ordering: runtime object submitted before the row exists, no worker queued, early reads stay Creating (test_synchronous_create_submits_runtime_before_persisting)
    • failed submit → SNAPSHOT::RUNTIME_CREATE_FAILED, nothing persisted (test_synchronous_create_failure_returns_error_without_persisting)
    • persist failure after submit → best-effort runtime object cleanup (test_synchronous_create_cleans_up_runtime_object_when_persist_fails)
    • recovery retries k8s/fsb inline without the pool (test_recovery_retries_synchronous_create_inline_without_worker)
    • HA recovery with fsb: peer replica resubmits the idempotent intent and never maps the row to Failed (unit: test_ha_recovery_converges_fsb_snapshot_without_failed_mapping; integration: test_fsb_snapshot_row_is_recovered_by_peer_replica)
    • creator-crash HA test updated to the synchronous flow; docker keeps worker-based coverage

Breaking Changes

  • None
  • Yes (describe impact and migration path)

Behavior note (not a contract change): for k8s/fsb, a runtime submit failure now surfaces as a synchronous 500 on POST instead of eventually transitioning an accepted row to Failed. The 202 + Creating + poll contract is unchanged.

Checklist

  • Linked Issue or clearly described motivation
  • Added/updated docs (if needed)
  • Added/updated tests (if needed)
  • Security impact considered
  • Backward compatibility considered

…untimes

Root fix for the early-read Failed race (opensandbox-group#2092, follow-up to opensandbox-group#2077).

k8s and fsb create_snapshot are fast, idempotent submits (CR name / fsb
request_id derive from the snapshot id), so POST /snapshots now creates
the runtime object before persisting the Creating row and returns without
touching the worker pool. Once a row is readable, the runtime object
always exists, so read-time sync and recovery can never map a
not-yet-created object to Failed, including on other HA replicas. A
persist failure after runtime creation triggers best-effort deletion of
the runtime object. Recovery retries stuck k8s/fsb rows inline instead
of resubmitting workers, keeping the pool and the in-flight guard
docker-only. Docker (slow container.commit) keeps the existing async
create path unchanged.
@github-actions github-actions Bot added component/server documentation Improvements or additions to documentation size/L Denotes a PR that changes 100-499 lines, ignoring generated files. labels Sep 30, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

component/server documentation Improvements or additions to documentation size/L Denotes a PR that changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

refactor(server): make snapshot create path synchronous for k8s/fsb runtimes (root fix for early-read Failed)

1 participant