Repository navigation
docs(rfc): explainable PreparedContext receipts - #1435
Tsukikage7 wants to merge 5 commits into
Conversation
Add bilingual RFCs for an opt-in, ephemeral PreparedContext Receipt on prepare, without changing default injection or host recall.
Match the RFC filename and navigation entry to the pull request number.
|
resolve conflicts |
1 similar comment
|
resolve conflicts |
|
please update & resolve conflicts |
Teingi
left a comment
There was a problem hiding this comment.
Reviewed against current master 86537e4f. An opt-in Receipt attached to the same prepare response is still useful: the public response still lacks an exact selection record. Keeping it ephemeral and reusing existing exact reads and authorization remains appropriate, with no new Receipt tables or database migration needed. Existing omission counters and recall-effort diagnostics should also be reused.
The inline comments identify four contract gaps that need to be resolved before implementation. This design is outdated and needs to be updated to reflect the current implementation.
|
|
||
| | Field | Contract | | ||
| | --- | --- | | ||
| | `kind` | `memory` or `experience` | |
There was a problem hiding this comment.
[P2] Cover the origins that prepare already supports
Current prepare can include Topic Memory by default, Profile through assembly, and code through include_code. Code references live in PreparedContextBuild.code_origins, so a code-only result can be ready with nonempty content and an empty origins tuple. Restricting kind to Memory/Experience and matching only origins therefore cannot describe all injected content; normal supported requests would either lose references or lose the Receipt entirely. The fixed 16-item bound also conflicts with supported configurations: context_assembly_max_entries=18 permits an assembly of 8 Memory, 8 Topic Memory, and 2 Experience items. Please update the selected-item types, completeness invariant, and bounds to cover the current prepare contract.
There was a problem hiding this comment.
Updated the selected-item contract in ca60d3d to cover Memory, Topic Memory, Experience, Profile, and code. Completeness now requires ordered identities from origins followed by code_origins, including code-only results and final truncated code ranges/hashes. The cap is 32, covering up to 26 historical items plus four code items; the 8 KiB bound drops the whole Receipt on overflow rather than truncating identities, and requires representative payload-size measurements before implementation ships.
This revises the bilingual RFC only; Runtime implementation remains deferred until the design is accepted. Local make docs-test and uv run --locked prek run -a passed; remote CI has not yet been confirmed passing for this head.
| | Field | Contract | | ||
| | --- | --- | | ||
| | `kind` | `memory` or `experience` | | ||
| | `memory_citation` or `artifact_ref` | Exact identity already admitted by the Builder | |
There was a problem hiding this comment.
[P2] Preserve the owning Scope in every selected identity
MemoryCitation and ArtifactRef alone do not identify the owning Scope. Current prepare merges the current Scope with its Context References and preserves cross-Scope origins as MemoryEntryAddress or ArtifactAddress. The database keys include Scope, so two Scopes can legitimately contain the same Artifact/revision/entry identifiers. Dropping that address makes progressive inspection ambiguous or causes it to read the wrong Scope; keeping the stated exact-origin invariant would instead reject these Receipts. Please retain the complete Scope-qualified identity and expand it through the existing read and authorization path for that Scope. This requires no new Receipt table.
There was a problem hiding this comment.
Updated in ca60d3d: every selected identity is Scope-qualified, including refs from the current Scope. Memory uses the existing MemoryEntryAddress shape, other Artifact families use ArtifactAddress, and code retains CodeEvidenceRef.scope_id. Identical short IDs in different Scopes remain distinct; ambiguous identities invalidate the Receipt. Progressive inspection reuses existing exact-read APIs, with no Receipt table or content cache.
This is an RFC contract revision, not an implemented Runtime change; cross-Scope ID collisions and exact-read reuse are included in implementation acceptance.
| used_bytes: integer, equal to content_bytes | ||
| truncated: true when any selected item was size-truncated | ||
| retrieval: | ||
| memory_mode: auto | fts | vector | hybrid | none |
There was a problem hiding this comment.
[P2] Define retrieval evidence for multiple Scope searches
One prepare request already searches each participating Scope separately. With complete vectors in the current Scope and incomplete vectors in a referenced Scope, auto resolves to hybrid for one search and fts for the other, and both can contribute injected items. A single memory_mode cannot represent that actual path; reporting auto would report the requested policy rather than the executed modes. The recall gate can also perform multiple rounds. Please define bounded evidence or aggregation semantics that preserve the actual modes and fallback outcomes across those searches, instead of assuming one Memory search result per prepare.
There was a problem hiding this comment.
Replaced the single memory_mode with bounded, per-round evidence in ca60d3d. Searches aggregate by round, family, actual mode, outcome, and fallback reason, preserving simultaneous hybrid/FTS paths without reporting auto as an executed mode. Normal FTS selection, inference fallback, and reuse of an earlier FTS fallback remain distinct. Round metadata identifies retained versus abandoned pools when expansion fails and returns round-zero candidates; executed searches from abandoned rounds remain visible.
The bounds allow three rounds and 32 search groups, with whole-Receipt omission on overflow. These are bilingual RFC requirements; request-local execution metadata still needs to be carried through the Runtime in the implementation PR.
| omitted. Stage names match the existing Runtime span names so a Receipt can be correlated with a trace without copying | ||
| span payloads. | ||
|
|
||
| Rerank: when Memory search produced a `MemoryRerankTrace`, the Receipt copies `policy_id` and `used_fallback` only. It |
There was a problem hiding this comment.
[P2] Retain the effective rerank configuration and non-determinism evidence
LLMMemoryReranker.policy_id is a fixed instruction-version string, while the current runtime allows a separate rerank model, configurable model settings, and Scope-specific memory.rerank Prompt revisions. Copying only that ID and the fallback flag therefore makes materially different model-backed executions indistinguishable. This does not meet #1356's explicit requirement for exact configuration evidence and a non_deterministic marker. Please define content-free fields for the effective model, safe configuration identity or digest, effective Prompt revision or digest, and non-deterministic status before freezing the Receipt contract.
There was a problem hiding this comment.
Added effective rerank configuration evidence in ca60d3d: credential-free provider/model identity, safe effective settings after inheritance/overrides, timeout/request limits, and a versioned canonical configuration digest. The Prompt identity comes from the ResolvedPrompt bound to the invocation and includes Scope, exact custom revision or built-in version, and compiled_digest; it cannot be replaced by the latest Prompt head or a fixed instruction policy ID. Repeated configurations are deduplicated and referenced by round/outcome groups.
Model-backed rerank is explicitly non_deterministic=true, including temperature-zero and fallback executions. Missing or unsafe configuration evidence omits the whole Receipt. This updates the RFC contract and acceptance criteria only; Runtime/OpenAPI implementation remains deferred until acceptance.
| offer a Server deployment default that turns Receipts on for every prepare. | ||
|
|
||
| `query`, `scope_id`, and `max_bytes` keep their current bounds. `query_digest` hashes the same normalized query Memory | ||
| search already uses (`normalize_text`: NFC Unicode of the trimmed query, UTF-8). The Receipt stores only `sha256:<hex>`. |
There was a problem hiding this comment.
[P2] query_digest names the wrong normalizer: Memory search uses normalize_query, and the two disagree inside the contract's own bound
The Receipt is specified to hash "the same normalized query Memory search already uses (normalize_text: NFC Unicode of the trimmed query, UTF-8)". Memory search does not use normalize_text — it uses normalize_query:
-
src/powercontext/builtin/artifacts/memory/service.py:717—normalized_query = normalize_query(query)is the retrieval path;normalize_textat:1720/:1827is the durable-entry-body path. -
The two functions differ exactly where it matters.
normalize_queryhas this docstring (canonical.py:90-97):A query is never persisted, so the durable entry body's UTF-8 byte bound does not apply to it. Callers that expose a query length limit enforce it on their own terms: the HTTP contract states its bound in characters, and applying the storage-layer byte bound here would reject queries the contract accepts.
normalize_texthard-rejects anything over 8192 UTF-8 bytes (canonical.py:85-86).
The query bound is expressed in characters on both sides of the contract — openapi/powercontext.yaml's PrepareContextRequest.query is maxLength: 8192, and models.py:233 is max_length=8192. So the two normalizers diverge on inputs the contract accepts. Measured against the head tree:
2730 汉字 = 8190 UTF-8 字节 | normalize_query=OK | normalize_text=OK
2731 汉字 = 8193 UTF-8 字节 | normalize_query=OK | normalize_text=REJECT(memory entry text must not exceed 8192 UTF-8 bytes)
3000 汉字 = 9000 UTF-8 字节 | normalize_query=OK | normalize_text=REJECT(memory entry text must not exceed 8192 UTF-8 bytes)
2731 characters is a legal query. Implementing the RFC literally would hash with the stricter function, and for a query past the byte bound it would raise inside Receipt assembly — landing on this RFC's own fail-open path, so a valid request silently receives no Receipt and the recorded reason is a local normalization limit rather than anything about method selection.
Suggested direction: name normalize_query here (and in zh:116), and drop the parenthetical's "UTF-8" framing so it does not imply the byte bound applies. If the intent was actually to state the hashing algorithm rather than the function, say that explicitly instead of naming a function that retrieval does not use.
Test worth adding when the Receipt is implemented: a non-ASCII query of at least 2731 characters produces the same query_digest as Memory search does for the same query.
Probe: probes/normalize_query_vs_text_probe.py.
| # The transaction has rolled back. Connection-level restoration must | ||
| # also succeed before the enclosing context can return this expiry. | ||
| sqlite_code = getattr(error.orig, "sqlite_errorcode", None) | ||
| if not committing and sqlite_code == SQLITE_INTERRUPT and loop.time() >= deadline: |
There was a problem hiding this comment.
[P2] Turning an interrupted body into a retryable attempt lets one record hold the write lock for its whole budget
The new guard here is what makes an interrupted body retryable:
if not committing and sqlite_code == SQLITE_INTERRUPT and loop.time() >= deadline:
raise ModelUsageAttemptExpired from errorThat is the right behaviour — a native interrupt with an unknown commit outcome should be retried rather than dropped, and the not committing guard keeps a COMMIT whose reply was lost from being counted twice. The consequence for how long the lock is held is worth stating, because it changes the bound this repository documents elsewhere.
The recorder slices its budget: _WRITE_ATTEMPT_SLICES = 4 and attempt_timeout = max(write_timeout_seconds / 4, 0.02) (src/powercontext/builtin/runtime/_model_usage.py:45, :219). On the first attempt allowance = attempt_timeout; after any repeat first_attempt = False and allowance = remaining, then _model_usage_transaction(min(remaining, allowance)) (:230, :232). So a retry runs inside the record's entire remaining budget, not one slice of it.
Before this change the same interruption produced an OperationalError that is_transaction_contention did not classify as contention, so the record was dropped and the lock released immediately. Now the retry path can occupy the write lock for up to model_usage_write_timeout_seconds in total rather than a quarter of it. tests/e2e/test_topic_memory_generic_api.py:451 and :582 pass model_usage_write_timeout_seconds=30.0, and runtime/config.py:230 caps that field at 30, so the upper bound there moves from ~7.5s to ~30s.
The trade is worth it — a slow body no longer loses records it still owns budget for. Two things would make it legible:
- State the new bound where the recorder decides it, so nobody has to rediscover that
write_timeout_secondsis now a per-record ceiling rather than a per-attempt one. - If blocking a business write for the full budget is not acceptable, cap the retry allowance separately (for example at
max(write_timeout / 4, 0.02)) instead of handing itremaining.
I could not measure the wall-clock difference reliably on this machine — the interrupt path did not fire in my runs of the recorder, so I am not quoting a blocked-write duration. The code path above is what I verified; the numbers I have seen elsewhere for it should be re-measured before being treated as settled.
The MySQL/OceanBase path is untouched by this PR, and I did not run it — no OceanBase available locally.
Which issue or RFC does this PR close?
Closes #1356.
Rationale for this change
POST /v1/context/preparealready selects, cites, and renders bounded context, but the public response is onlystatus,content, andcontent_bytes. Exact origins stay inside the Runtime and are discarded at the HTTP boundary. Operators, hosts, and evaluation therefore cannot tell what was selected, why items were omitted, or which retrieval path produced the injected bytes.Native SQLite usage deadlines can also interrupt an uncommitted transaction body and drop an accepted record despite remaining write budget. The usage fix retries only these expired, rolled-back bodies; an interrupted COMMIT with an unknown outcome remains non-retryable.
What changes are included in this PR?
include_receiptfield on the existing prepare operation.Are there any user-facing changes?
The Receipt remains a design proposal; the optional prepare field and Receipt payload require a later implementation PR after acceptance. The existing usage-accounting fix preserves accepted increments when a native SQLite deadline interrupts the body and budget remains. It adds no public API, schema migration, larger write budget, or retry of indeterminate commits.
How was this change tested?
make docs-test: website lint, tests, link validation, production build, and static export verification passed.Verified both English and Chinese RFC pages in the static export; export verification checked 877 public pages and their internal links.
uv run --locked prek run -a: all checks passed, including type checking.git diff upstream/master...HEAD --check.Checked relative RFC links, JSON examples, and matching contract fields in both locales.
Python 3.12 focused validation:
python -m pytest tests/builtin/runtime/test_model_usage_recorder.py tests/e2e/test_topic_memory_generic_api.py -q— 50 passed, including the unchanged HTTP cancellation scenario.Real SQLite reproduction before/after the fix: first-attempt native SQL interruption produced code 9 and zero persisted requests before the fix; the fixed path retried and persisted exactly one request. The reproduction establishes this defect; the original CI warning hides the exception, so it does not prove this was that run's sole cause.
Python 3.12.14 full unit suite:
python -m pytest --doctest-modules --ignore=tests/e2e— 3374 passed, 72 skipped. Full E2E:python -m pytest tests/e2e— 330 passed, 65 skipped, 4 deselected. Both exited 0 at head6da77b1e.Local macOS validation used a fresh temporary virtualenv because Python 3.12.14 ignores hidden editable
.pthfiles in the managed worktree environment. Unit tests isolatedDSH_HOMEand global/system Git configuration to exclude local-user configuration; user files were unchanged.Remote workflows for head
6da77b1erequire maintainer approval (action_required); no passing remote CI is claimed for this head.AI usage statement
OpenAI Codex assisted with the bilingual Receipt contract, the existing SQLite usage-accounting fix and regression tests, and documentation/repository validation. A specific implementation-model identifier was not recorded.