Skip to content

feat(memory): gate writes on decision-model evidence sufficiency - #1742

Merged
Teingi merged 17 commits into
oceanbase:masterfrom
AlexStocks:feat/memory-write-gate
Sep 29, 2026
Merged

Teingi merged 17 commits into
oceanbase:masterfrom
AlexStocks:feat/memory-write-gate

Conversation

@AlexStocks

@AlexStocks AlexStocks commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #1740. This branch is based on that PR's head, so the diff below currently also
contains #1740's changes. Please review #1740 first; once it lands, this PR's diff shrinks to the
files listed at the bottom.

Related to #1644. That issue asks whether a decision model can judge semantic relevance and
evidence coverage. This PR implements the consumer for the evidence-sufficiency half, behind a
default-off switch. It closes nothing.

Rationale for this change

#1740 added a decision role to the Runtime but deliberately shipped it with no consumers. This PR
adds the first one: a write-time evidence gate for Memory.

When a Memory write arrives — from automatic extraction or from an explicit remember — the gate
asks one bounded question: does the cited evidence actually cover the claim being stored? Today
only the existence of the cited ids is checked; sufficiency is not checked at all. A model-backed
gate can judge that, at a fraction of the cost of routing the whole write through a general
generation model.

The verdict is advisory, and a refusal is visible

This is the part worth reviewing closely.

  • ACCEPT — the write proceeds unchanged.
  • FLAG — the write proceeds, annotated. The gate fills the change's reason only when the
    candidate did not already carry one, so it never overwrites an existing explanation.
  • HOLD — the write does not happen, and the caller is told why, through structured
    refusal rather than a log line:
    • automatic extraction: flush() returns held_count / hold_codes, and the window cursor
      still advances;
    • explicit remember / revise: MemoryWriteRejectedError is raised, carrying code and
      reason;
    • any plan_remember caller: reads MemoryWritePlan.decision directly.

A HOLD is not routed to the Review Inbox, and that is deliberate. For Memory, an unstored
observation is not a pending review item: regular Memory writes commit directly, and rejected
content is corrected through its existing revision semantics. The inverse of a silent drop is a
visible, reasoned refusal — not a queue. Reuse of needs_evidence and
evidence_limit_exceeded keeps the refusal vocabulary consistent with the existing evidence
resolution errors; insufficient_coverage is added for the case where evidence exists but does not
cover the claim.

Fail-open, and no thresholds

  • If the gate is unavailable — not configured, ABSTAIN, or used_fallback=true — the write
    proceeds. A gate that cannot judge must never be the reason a memory is lost.
  • The gate is disabled by default and makes no model call until it is switched on.
  • Enabling the gate while no decision backend is available does not fail startup. This is
    deliberately the opposite of the decision role itself, which fails fast when it is explicitly
    enabled with an unusable backend: the gate is auxiliary and fail-open by contract, so it logs a
    structured warning naming the event and passes writes through. A misconfigured gate must never
    block Memory writes — but it must never do so silently either.
  • A gate passed in explicitly takes precedence over one built from configuration.
  • One combined configuration is worth naming: if the decision role itself is enabled with no usable
    backend, startup fails fast by design, and the gate's pass-through warning is emitted first. The
    failure still belongs to the decision role and the outcome is correct; only the order of those two
    log lines is arbitrary.
  • No threshold values are shipped. The configuration exposes a direction only
    (memory_write_gate_hold_on), with the threshold left unset. Which direction a backend's
    confidence actually points is a property of that backend and must be established by probing it
    with a known-answer pair (a real preference versus filler) before a value is trusted. The
    contract tests carry that probe, including a case where a reverse-polarity backend makes the
    probe fail.

What changes are included in this PR?

  • runtime/memory_write_gate.py — the decision-backed gate adapter, plus DecisionKind values for
    the two consumers.
  • artifacts/memory/protocols.py — the gate's value objects and port (MemoryWriteGate,
    MemoryWriteGateRequest, MemoryWriteAssessment, MemoryWriteVerdict,
    MemoryWriteRejectionCode). They live in the Memory family rather than in runtime/ because the
    Runtime package eagerly imports relational → memory.service → protocols; putting them in
    runtime/ would force a reverse import and create a cycle. The dependency direction stays
    one-way: runtime → memory. All public names remain importable from
    powercontext.builtin.runtime.memory_write_gate.
  • artifacts/memory/service.py — assessment between candidate selection and commit preparation.
  • artifacts/memory/errors.py — MemoryWriteRejectedError(code, reason).
  • runtime/relational.py, runtime/models.py — a held write does not store Memory, does not
    interrupt the window, and is reported through held_count / hold_codes.
  • runtime/application.py — explicit writes surface the refusal through the existing error
    channel.
  • runtime/composition.py — resolves the gate for the public runtime API: an explicit injection
    wins, configuration otherwise builds one, and an enabled gate with no usable backend logs a
    structured warning and passes writes through rather than failing startup.
  • runtime/config.py, .env.example — the gate's settings.
  • Three test modules.

Are there any user-facing changes?

Yes, but only when the gate is switched on:

  • Three new settings; all inert by default, and all three are read by production code — this PR
    deliberately adds no setting that nothing consumes.
  • MemoryFlushResult gains held_count and hold_codes (defaulted, so existing readers are
    unaffected), and MemoryWriteRejectedError is a new exception which callers of explicit writes
    may now receive.

With the gate off, nothing is constructed, no model call is made, and every write behaves exactly
as before. When on, a refused write surfaces to the caller; it is never silently discarded.

No new dependency; pyproject.toml and uv.lock are untouched. No persistence schema change.

How was this change tested?

  • ruff check . — All checks passed
  • ruff format --check . — 890 files already formatted
  • ty check --python-version 3.11 <changed src files> — All checks passed
  • Global ty check — no diagnostic points at any file in this PR
  • The modules this PR adds or changes, plus the HTTP API contract — 78 passed
  • pytest tests/builtin/runtime — 478 passed, 7 failed on the development host

The seven failures are all tests that spawn a real child process and then time out. They are
attributed to host load rather than to this change: one of them fails in isolation as well, none of
the affected modules reference the gate, and 478 + 7 = 485 matches the count from an earlier green
run of the same suite on the same code path. CI runs this suite on four Python versions, and every
tests (3.11) through tests (3.14) job is green on this commit, together with quality,
windows-unit-portability and both Acceptance jobs.

The behaviour was exercised through all four paths — accept, flag, hold, and fail-open — including
asserting that a held automatic write advances the window without storing Memory, that a held
explicit write raises with code and reason readable by the caller, and that an unavailable gate
lets the write through.

Two of these assertions were additionally validated by mutation, so that they are load-bearing
rather than incidental:

  • Removing the guard that preserves an existing candidate reason makes the flagged-write
    regression fail with 'evidence is thin' != 'an explicit reason'.
  • Renaming the structured event on the pass-through warning makes the unavailable-backend test
    fail, reporting the event set it actually observed.

AI usage statement

AI assistance was used to develop this change: the implementation and its tests were produced with
a WorkBuddy agent (Claude-family model) working from a human-reviewed design. The diff, and every
command listed above, were reviewed and run by the author before opening this PR.

Add the cross-family decision role contract (DecisionModel, DecisionRequest, DecisionResult, DecisionOutcome, DecisionInput/Output, LLMDecisionModel, FailOpenDecisionModel) with its configuration surface, composition wiring, and a non-blocking inference.decision readiness probe.

The role is disabled by default and has no consumers: no decision backend is built and no model call is made until runtime.decision_assistance_enabled is set or a backend is injected. The shared FailOpenDecisionModel envelope turns any backend failure into a no-op abstention while asyncio.CancelledError propagates, so callers need no try/except at the call site.

Decision support depends on no persistence schema: it adds no tables, migrations, or processing capabilities, and leaves canonical_processing_manifest unchanged. It is never registered as an MCP tool.
The decision-seam tests must pass the repository-wide 'ty check' (make check runs it globally, including tests/), not only the five production files.

- Type the _config(**runtime) and InferenceConfig(**overrides) helpers with Any so per-field splats are accepted.
- Build decision_base_url through AnyHttpUrl, matching the existing inference-endpoint tests.
- Suppress the intentional frozen-dataclass assignment with ty's own '# ty: ignore[invalid-assignment]' rule code, which ty recognizes (the previous mypy '# type: ignore[misc]' did not apply).
- Narrow the Memory | None returned by remember() before passing it to search().
The `memory_write_gate_enabled` switch was inert: `open_builtin_runtime`
only offered a pass-through injection point, so enabling the gate through
configuration left `MemoryService._write_gate` as `None` and silently
accepted every write. The composition root now builds the gate from
configuration via `build_memory_write_gate`, keeping an explicit
injection authoritative over the configuration-derived one.

The gate stays auxiliary and fail-open, unlike the fail-fast decision
role: when it is enabled but no decision backend is available, the runtime
logs a warning (`memory.write-gate.unavailable`) and passes writes through
instead of failing startup, so a misconfigured gate can never block a
Memory write.

Also locks a regression: a FLAG verdict must preserve an existing
candidate reason rather than overwrite it.
`handoff_escalation_enabled` was introduced by 10b9505 with zero
production consumers: only a declaration, two assertions that the field
exists, and a stale .env.example line. Handoff consult escalation is a
later batch and no reader was wired, so shipping the switch as-is would
re-introduce the exact "configured but never read" silent failure this
batch exists to remove.

- remove the field from RuntimeConfig
- remove its two test assertions and the field-only test function
- remove the stale POWERCONTEXT_SERVER_RUNTIME_HANDOFF_ESCALATION_ENABLED
  documentation line

Also give the gate-unavailable test teeth: assert the structured
`event == "memory.write-gate.unavailable"` record (not just the
human-readable message), proven by mutating the production event string
and observing the test fail. Read the extra attribute via getattr to
match the repo's caplog idiom and stay clean under the global type check.
Propagate Memory write gate holds through service, runtime, flush transport, and HTTP error mapping. Feed bounded source content to the gate, fail open on injected gate failures, and hold oversized candidate batches instead of accepting unassessed tails.

Tested: uv run --no-sync pytest tests/builtin/runtime/test_memory_write_gate_contract.py tests/builtin/runtime/test_memory_write_gate_paths.py -q; uv run --no-sync pytest tests/test_server_generation.py -q; ruff check/format --check changed files; ty check changed files
Synchronize checked-in generated HTTP schema and models after adding Memory write gate hold details to the OpenAPI contract.

Tested: uv run --no-sync python scripts/generate_api.py --check; uv run --no-sync python scripts/generate_js_operations.py --check; uv run --no-sync pytest tests/test_api_contract.py tests/test_js_operations.py -q; uv run --no-sync pytest tests/builtin/runtime/test_memory_write_gate_contract.py tests/builtin/runtime/test_memory_write_gate_paths.py -q; ruff/ty on generated files
@AlexStocks

AlexStocks commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor Author

@Teingi I have handled all your review comments. please review this pr again.

AlexStocks and others added 3 commits September 27, 2026 21:04
# Conflicts:
#	src/powercontext/builtin/runtime/__init__.py
#	src/powercontext/builtin/runtime/config.py
#	src/powercontext/builtin/runtime/decision_model.py
#	tests/builtin/runtime/test_decision_model.py
Lore: PR oceanbase#1742 must track upstream master after decision timeout and LongMemEval v2 landed, while preserving the memory write gate runtime injection chain.

Constraint: keep master decision timeout behavior and explicit decision model contracts while retaining the oceanbase#1742 write-gate configuration path.

Scope-risk: narrow; conflict resolution touched decision runtime composition and its focused tests.

Tested: uv run --no-sync pytest tests/builtin/runtime/test_decision_model.py tests/builtin/runtime/test_decision_composition.py tests/builtin/runtime/test_decision_config.py tests/builtin/runtime/test_decision_default_off.py tests/builtin/runtime/test_decision_fail_open.py tests/builtin/runtime/test_decision_schema_decoupled.py tests/builtin/runtime/test_memory_write_gate_contract.py tests/builtin/runtime/test_memory_write_gate_paths.py tests/test_server_generation.py -q

Tested: uv run --no-sync ruff check src/powercontext/builtin/runtime/composition.py src/powercontext/builtin/runtime/config.py src/powercontext/builtin/runtime/decision_model.py src/powercontext/builtin/runtime/relational.py tests/builtin/runtime/test_decision_composition.py

Tested: uv run --no-sync ruff format --check src/powercontext/builtin/runtime/composition.py src/powercontext/builtin/runtime/config.py src/powercontext/builtin/runtime/decision_model.py src/powercontext/builtin/runtime/relational.py tests/builtin/runtime/test_decision_composition.py

Tested: uv run --no-sync ty check --python-version 3.11 src/powercontext/builtin/runtime/composition.py src/powercontext/builtin/runtime/config.py src/powercontext/builtin/runtime/decision_model.py src/powercontext/builtin/runtime/relational.py tests/builtin/runtime/test_decision_composition.py

Co-authored-by: OmX <omx@oh-my-codex.dev>
Lore: PR oceanbase#1742 review found that Memory write gate assessment pooled window evidence, lost candidate citation mapping, missed spawned worker gate construction, and made early budget holds hard to observe.

Constraint: judge only each candidate's effective canonical citations, keep backend failures fail-open, and do not apply semantic verdicts to omitted or unmaterialized evidence.

Scope-risk: focused on Memory write gate evidence projection, scheduled hold observability, and child worker runtime composition.

Tested: uv run --no-sync pytest tests/builtin/runtime/test_memory_write_gate_contract.py tests/builtin/runtime/test_memory_write_gate_paths.py tests/builtin/runtime/test_family_processing.py::test_spawned_memory_worker_reconstructs_configured_write_gate tests/builtin/runtime/test_decision_composition.py tests/builtin/runtime/test_decision_config.py tests/builtin/runtime/test_decision_default_off.py tests/builtin/runtime/test_decision_fail_open.py tests/builtin/runtime/test_decision_schema_decoupled.py tests/test_server_generation.py -q

Tested: uv run --no-sync ruff check src/powercontext/builtin/artifacts/memory/service.py src/powercontext/builtin/runtime/application.py src/powercontext/builtin/runtime/composition.py src/powercontext/builtin/runtime/family_processing.py tests/builtin/runtime/test_memory_write_gate_paths.py tests/builtin/runtime/test_family_processing.py

Tested: uv run --no-sync ruff format --check src/powercontext/builtin/artifacts/memory/service.py src/powercontext/builtin/runtime/application.py src/powercontext/builtin/runtime/composition.py src/powercontext/builtin/runtime/family_processing.py tests/builtin/runtime/test_memory_write_gate_paths.py tests/builtin/runtime/test_family_processing.py

Tested: uv run --no-sync ty check --python-version 3.11 src/powercontext/builtin/artifacts/memory/service.py src/powercontext/builtin/runtime/application.py src/powercontext/builtin/runtime/composition.py src/powercontext/builtin/runtime/family_processing.py tests/builtin/runtime/test_memory_write_gate_paths.py tests/builtin/runtime/test_family_processing.py

Co-authored-by: OmX <omx@oh-my-codex.dev>

@AsperforMias AsperforMias left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Reviewed the current head, including gate assembly in background workers, per-candidate evidence resolution and budgets, and visible refusal through explicit writes and automatic processing. No blocking issue found. The model-quality experiments in #1644 remain follow-up scope, as stated in the PR.

Resolve the adjacent-addition conflicts with the memory capacity contract
(oceanbase#1746) by keeping both seams:

- RuntimeConfig keeps the write-gate fields and the capacity/compaction
  budgets, including validate_memory_capacity_order.
- MemoryService and the scope relational wiring keep the write gate plus
  the capacity budget, compaction policy, and history limit.
- The Server error mapping keeps the extracted conflict helper (it already
  owns revision_conflict and memory_entry_inactive) and adds the
  memory_write_rejected branch on top.
Raise the hermetic Topic Memory worker timeout so the spawned worker can cold-start under Python 3.13 and still publish before the bounded search deadline.

Capture the current artifact_processing.failed supervisor event in R8 diagnostics so future worker failures keep their redacted classification.

Tested: WSL SETUPTOOLS_SCM_PRETEND_VERSION=1.1.1.dev0 UV_PROJECT_ENVIRONMENT=.venv-linux-313 uv run --python 3.13 pytest tests/e2e/test_topic_memory_product_chain.py::test_r8_e0_runs_the_complete_hermetic_topic_product_chain tests/e2e/test_topic_memory_product_chain.py::test_worker_failure_capture_records_current_supervisor_event -q

Tested: uv run --no-sync ruff check tests/e2e/topic_memory_product/common.py tests/e2e/topic_memory_product/harness.py tests/e2e/test_topic_memory_product_chain.py

Tested: uv run --no-sync ruff format --check tests/e2e/topic_memory_product/common.py tests/e2e/topic_memory_product/harness.py tests/e2e/test_topic_memory_product_chain.py

Tested: uv run --no-sync ty check --python-version 3.11 tests/e2e/topic_memory_product/common.py tests/e2e/topic_memory_product/harness.py tests/e2e/test_topic_memory_product_chain.py
@AlexStocks

Copy link
Copy Markdown
Contributor Author

@Teingi please review it again.

@AlexStocks

Copy link
Copy Markdown
Contributor Author

Pushed b629207b to fix the latest CI failures.

Root cause: the R8 hermetic Topic Memory test used a 20s spawned worker timeout, which was too tight for Python 3.13/Ubuntu cold startup before topic generation. The worker could time out before publishing a searchable topic. I raised that bounded test timeout to 60s and also updated the R8 worker-failure capture to recognize the current artifact_processing.failed supervisor event so future failures keep their redacted classification.

Validation:

  • WSL SETUPTOOLS_SCM_PRETEND_VERSION=1.1.1.dev0 UV_PROJECT_ENVIRONMENT=.venv-linux-313 uv run --python 3.13 pytest tests/e2e/test_topic_memory_product_chain.py::test_r8_e0_runs_the_complete_hermetic_topic_product_chain tests/e2e/test_topic_memory_product_chain.py::test_worker_failure_capture_records_current_supervisor_event -q (2 passed)
  • uv run --no-sync ruff check tests/e2e/topic_memory_product/common.py tests/e2e/topic_memory_product/harness.py tests/e2e/test_topic_memory_product_chain.py
  • uv run --no-sync ruff format --check tests/e2e/topic_memory_product/common.py tests/e2e/topic_memory_product/harness.py tests/e2e/test_topic_memory_product_chain.py
  • uv run --no-sync ty check --python-version 3.11 tests/e2e/topic_memory_product/common.py tests/e2e/topic_memory_product/harness.py tests/e2e/test_topic_memory_product_chain.py
  • GitHub checks are now green, including the previously failing tests (3.13) and Acceptance (sqlite).

Encode Memory write-gate candidates with explicit indices so multiline batches cannot collide, and read registered text-evidence projections for local Sources before marking evidence incomplete. Keep explicit Artifact families authoritative when recovering generic Artifact candidates.

Surface spawned Memory HOLD outcomes through worker completions and parent logs, retain hold fields in JSON logging, and reject injected Memory gates when spawned workers cannot reconstruct them.

Constraint: preserve existing window consumption semantics for held writes while making the refusal observable.
Scope-risk: limited to Memory write-gate input construction, processing-worker completion metadata, and operational logging fields.
Tested: .\.venv\Scripts\python.exe -m pytest tests/builtin/runtime/test_memory_write_gate_contract.py::test_candidate_subject_preserves_candidate_boundaries tests/test_server_logging.py::test_json_formatter_emits_stable_operational_fields -q
Tested: uv run --no-sync ruff check src/powercontext/builtin/runtime/memory_write_gate.py src/powercontext/builtin/artifacts/memory/service.py src/powercontext/builtin/runtime/processing_contracts.py src/powercontext/builtin/runtime/family_processing.py src/powercontext/builtin/runtime/artifact_processing.py src/powercontext/builtin/runtime/composition.py src/powercontext/builtin/runtime/relational.py src/powercontext/server/logging.py tests/builtin/runtime/test_memory_write_gate_contract.py tests/builtin/runtime/test_memory_write_gate_paths.py tests/builtin/runtime/test_processing_composition.py tests/builtin/runtime/test_family_processing.py tests/test_server_logging.py
Tested: uv run --no-sync ruff format --check src/powercontext/builtin/runtime/memory_write_gate.py src/powercontext/builtin/artifacts/memory/service.py src/powercontext/builtin/runtime/processing_contracts.py src/powercontext/builtin/runtime/family_processing.py src/powercontext/builtin/runtime/artifact_processing.py src/powercontext/builtin/runtime/composition.py src/powercontext/builtin/runtime/relational.py src/powercontext/server/logging.py tests/builtin/runtime/test_memory_write_gate_contract.py tests/builtin/runtime/test_memory_write_gate_paths.py tests/builtin/runtime/test_processing_composition.py tests/builtin/runtime/test_family_processing.py tests/test_server_logging.py
Tested: uv run --no-sync ty check --python-version 3.11 src/powercontext/builtin/runtime/memory_write_gate.py src/powercontext/builtin/artifacts/memory/service.py src/powercontext/builtin/runtime/processing_contracts.py src/powercontext/builtin/runtime/family_processing.py src/powercontext/builtin/runtime/artifact_processing.py src/powercontext/builtin/runtime/composition.py src/powercontext/builtin/runtime/relational.py src/powercontext/server/logging.py tests/builtin/runtime/test_memory_write_gate_contract.py tests/builtin/runtime/test_memory_write_gate_paths.py tests/builtin/runtime/test_processing_composition.py tests/builtin/runtime/test_family_processing.py tests/test_server_logging.py
Tested: git diff --check
@AlexStocks

Copy link
Copy Markdown
Contributor Author

Verified at b629207b: I have marked 12 fixed review threads as resolved, including the four confirmed fixes from the latest review. Thank you for addressing those.

Five P2 threads remain open. A fresh replay with isolated databases confirms:

  • Candidate-to-citation mapping: (A\nB, C) and (A, B\nC) with the same per-index citations still produce identical complete judge requests. The latter persists B+C citing only sourceC.
  • Refusal visibility: The spawned worker holds an over-budget candidate, advances cursor/ack, and reports succeeded without exposing the refusal to its parent. The timer's JSON output also drops refusal codes/counts.
  • Injected gate in workers: The parent holds a direct write, but the child ignores the injected gate and persists the contradictory candidate.
  • Local Source projections: A registered short text projection is rejected as evidence_limit_exceeded before the gate is called; the same write succeeds with the gate disabled.
  • Explicit Artifact family: An input retaining family='skill' is persisted as an experience reference even with the gate disabled.

The candidate-mapping and refusal-visibility fixes still have uncovered paths; the other three findings concern additional cases. Please address these five before merge. Their existing threads contain the reproduction details and required behavior.

This replay used fresh SQLite databases, actual spawned workers and a real scheduler timer with controlled gate/model responses. It did not rerun external-model or OceanBase validation.

done

@Teingi
Teingi merged commit 6e23756 into oceanbase:master Sep 29, 2026
24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants