Repository navigation
feat(runtime): add decision policy governance types - #1921
AlexStocks wants to merge 6 commits into
Conversation
Add the internal DecisionPolicy, DecisionAssessment, and DecisionObservation vocabulary from RFC 1770. Route the Memory write gate through the shared assessment mapping before preserving its existing ACCEPT/FLAG/HOLD contract. Constraint: no public HTTP/OpenAPI surface, persistence schema, or hosted-provider configuration is added. Rejected: copying Hermes implementation code or plugin structure; only the RFC-level governance semantics are mirrored. Scope-risk: limited to runtime decision vocabulary and Memory write gate assessment mapping. Tested: UV_PROJECT_ENVIRONMENT=.venv312 uv run --python 3.12 --no-sync pytest tests/builtin/runtime/test_decision_policy.py tests/builtin/runtime/test_memory_write_gate_contract.py::test_memory_write_decision_policy_keeps_fallback_unadjudicated tests/builtin/runtime/test_memory_write_gate_contract.py::test_memory_write_decision_policy_uses_configured_hold_direction -q Tested: UV_PROJECT_ENVIRONMENT=.venv312 uv run --python 3.12 --no-sync ruff check src/powercontext/builtin/runtime/decision_policy.py src/powercontext/builtin/runtime/memory_write_gate.py src/powercontext/builtin/runtime/__init__.py tests/builtin/runtime/test_decision_policy.py tests/builtin/runtime/test_memory_write_gate_contract.py Tested: UV_PROJECT_ENVIRONMENT=.venv312 uv run --python 3.12 --no-sync ruff format --check src/powercontext/builtin/runtime/decision_policy.py src/powercontext/builtin/runtime/memory_write_gate.py src/powercontext/builtin/runtime/__init__.py tests/builtin/runtime/test_decision_policy.py tests/builtin/runtime/test_memory_write_gate_contract.py Not-tested: full Memory write gate async contract on this Windows host hangs in CPython Proactor event-loop socketpair setup; default uv Python 3.14 currently fails tests/conftest.py typing import; ty check cannot resolve pydantic/pytest in this environment. Co-authored-by: OmX <omx@oh-my-codex.dev>
Apply the RFC 1770 policy mode, privacy-boundary, and observation rules to the existing Memory write gate. Default configuration now prevents content egress, while local-only evaluation keeps the established opt-in enforcement path. Constraint: no HTTP/OpenAPI contract, persistence schema, durable observation ledger, hosted sanitizer, or provider integration is added. Scope-risk: limited to Memory write-gate configuration, decisions, privacy-safe observation logs, and request attribution. Tested: uv run --no-sync pytest tests/builtin/runtime/test_memory_write_gate_contract.py -q (26 passed) Tested: uv run --no-sync pytest tests/builtin/runtime/test_memory_write_gate_contract.py tests/builtin/runtime/test_memory_write_gate_paths.py -q (46 passed) Tested: uv run --no-sync ruff check <11 changed files>; uv run --no-sync ruff format --check <11 changed files>; uv run --no-sync ty check --python-version 3.11 <11 changed files> Not-tested: a final Windows full-runtime retry was stopped after exceeding the prior 175s run time. The only earlier full-suite failure was reproduced against the exact b1d544b baseline in the spawned-worker diagnostics path, unrelated to these files.
hidb4ai
left a comment
There was a problem hiding this comment.
Reviewed ba7c086. Three issues need addressing before merge: shadow/advisory can suppress Memory creation, the declared privacy boundary does not constrain the actual backend call, and the Server logger drops the new observation fields. Reproductions and required behavior are in the inline comments.
Persist a Scope-isolated, reference-only observation ledger for Memory write-gate decisions. Sidecar failures do not alter gate outcomes. Constraint: observations exclude candidate content, evidence bodies, model rationale, and exception messages. Scope-risk: retention prunes expired observations only within the Scope receiving a write. Tested: uv run --no-sync pytest tests/builtin/persistence/test_decision_observation_repository.py tests/builtin/runtime/test_decision_config.py tests/builtin/runtime/test_memory_write_gate_contract.py tests/builtin/runtime/test_memory_write_gate_paths.py; uv run --no-sync ruff check <16 changed files>; uv run --no-sync ruff format --check <16 changed files>; uv run --no-sync ty check --python-version 3.11 <16 changed files>.
Apply mode semantics to evidence preflight rejections, enforce loopback-only local evaluation, and make privacy outcomes observable in the private decision ledger and JSON logs. Constraint: observations retain bounded references and counts only; no candidate or evidence content is logged or stored. Scope-risk: unknown, LAN, wildcard, and hosted backends fail open under local_only until an explicitly loopback endpoint is configured. Tested: focused gate, policy, composition, logging, and MySQL schema suite (108 passed); Ruff, format, ty 3.11, and diff check.
|
resolve conflicts |
hidb4ai
left a comment
There was a problem hiding this comment.
Re-reviewed 62b4117 and replayed the original counterexamples: the shadow/advisory, privacy-boundary, and serialized-log findings are fixed. Two remaining issues in the new observation ledger need addressing; details and reproductions are inline.
| if not all(_SAFE_REFERENCE.fullmatch(value) for value in (*observation.subject_refs, *observation.evidence_refs)): | ||
| raise ValueError("decision observation references must be structured identifiers") # noqa: TRY003 |
There was a problem hiding this comment.
[P2] Preserve observations for supported Source identifiers
The reference allowlist is narrower than the supported Source ID contract. Through public capture/remember APIs with real SQLite, https://example.test/page, src/settings.py, and 会议记录 all write Memory successfully but leave zero ledger observations: this validation raises ValueError, which the sink caller catches and logs. The control ID safe persists one observation. This makes audit coverage depend on identifier spelling and loses normal URL, path, and Unicode references. Preserve supported references in a bounded structured representation or safe encoding instead of discarding the entire observation.
There was a problem hiding this comment.
Still reproduced on d4a8603 through public capture/remember APIs with real SQLite: safe writes Memory and one ledger observation, while https://example.test/page, src/settings.py, and 会议记录 each write Memory but persist zero observations. All three log decision.observation_sink_failed with ValueError. The current allowlist still rejects these supported Source IDs; the RFC and upstream-merge updates do not change this path. Please preserve these references in a bounded structured representation or safe encoding so the observation is retained.
| privacy_outcome=self._privacy_outcome(policy_assessment), | ||
| model_policy_id=None if policy_assessment.source.value == "none" else self.policy_id, | ||
| assessment=_safe_observation_assessment(policy_assessment), | ||
| final_action=f"memory_write_{assessment.verdict.value}", |
There was a problem hiding this comment.
[P2] Record the domain outcome separately from gate acceptance
This final_action is persisted before Memory commit preparation/application finishes. With memory_max_active_entries=1, a second remember() raises MemoryCapacityExceededError and leaves revision 1, yet the ledger records memory-write:memory@2 with final_action=memory_write_accept. An unapplied plan and a deduplicated no-op produce the same action. ACCEPT correctly describes gate permission, but it does not supply RFC 1770's required final domain action actually taken. Keep that assessment and record the actual outcome at the Memory owner boundary so failed, unapplied, unchanged, and applied operations remain distinguishable.
There was a problem hiding this comment.
Still reproduced on d4a8603: with memory_max_active_entries=1, the second remember() raises MemoryCapacityExceededError and Memory remains at revision 1, but the ledger records memory-write:memory@2 with final_action=memory_write_accept. Unapplied plans and deduplicated no-ops also record acceptance. The RFC update defers replay tooling but still requires the final domain action actually taken. This field is emitted before commit preparation/application and is never finalized; retain gate permission separately and record the actual outcome at the Memory owner boundary.
Clarify that replay and rescore are promotion evidence for enforcing policies, rather than mandatory runtime capabilities for every consumer. Track the Memory write-gate evidence work in #1930. Constraint: the RFC does not require additional consumers, a dashboard, replay tooling, or a cross-consumer evaluation bundle. Scope-risk: consumers without promotion evidence remain disabled, shadow, or advisory; existing authority, Scope, revision, and citation boundaries are unchanged. Tested: website pnpm lint; website pnpm test (13 passed); website pnpm build (255 public pages and 405 repository files link-validated; 950 public pages export-verified); git diff --check.
…cy-governance # Conflicts: # docs/en/docs/operate/troubleshoot.md # docs/zh/docs/operate/troubleshoot.md
There was a problem hiding this comment.
🟡 Changes recommended
Observation timing, reference encoding, policy consistency, and missing backend metadata currently violate the intended governance contract.
6 open findings
Observations record actions before they actually occur · New Effective policy diverges from operational gate behavior · New Recorded references cannot replay the evaluated candidate · New Valid identity references are rejected, losing durable observations · New Missing decision backend prevents default no-call gate observations · New Backend identity and latency are missing from observations · New
What changed in this PR
Adds governed decision-policy contracts and applies them to the Memory write gate with privacy modes, structured observations, and durable persistence.
Changes:
- Introduces decision policy, assessment, privacy, and observation types.
- Adds mode-aware Memory gating and scoped observation retention.
- Updates schema, recovery docs, logging, and focused tests.
| File | Description |
|---|---|
.env.example |
Documents gate modes and privacy settings. |
docs/en/docs/operate/troubleshoot.md |
Adds observation-table recovery guidance. |
docs/en/rfcs/1770-decision-policy-governance.md |
Clarifies promotion and replay requirements. |
docs/zh/docs/operate/troubleshoot.md |
Updates Chinese recovery guidance. |
docs/zh/rfcs/1770-decision-policy-governance.md |
Synchronizes RFC clarifications. |
src/powercontext/builtin/artifacts/memory/protocols.py |
Extends gate requests and preflight protocols. |
src/powercontext/builtin/artifacts/memory/service.py |
Supplies scoped references and observation sinks. |
src/powercontext/builtin/decision_observations.py |
Defines runtime-independent governance contracts. |
src/powercontext/builtin/persistence/__init__.py |
Exports the observation repository. |
src/powercontext/builtin/persistence/decision_observations.py |
Persists and validates observations. |
src/powercontext/builtin/persistence/tables.py |
Defines the observation ledger schema. |
src/powercontext/builtin/runtime/__init__.py |
Exports decision governance APIs. |
src/powercontext/builtin/runtime/composition.py |
Composes privacy-aware decision backends. |
src/powercontext/builtin/runtime/config.py |
Adds mode, privacy, and retention settings. |
src/powercontext/builtin/runtime/decision_model.py |
Propagates backend-locality attestations. |
src/powercontext/builtin/runtime/decision_observation_ledger.py |
Implements the relational observation sink. |
src/powercontext/builtin/runtime/decision_policy.py |
Maps results into governed assessments. |
src/powercontext/builtin/runtime/memory_write_gate.py |
Applies policy modes, privacy, and observations. |
src/powercontext/builtin/runtime/relational.py |
Wires scoped ledger persistence into Memory. |
src/powercontext/server/logging.py |
Adds decision fields to structured logs. |
tests/builtin/persistence/test_decision_observation_repository.py |
Tests ledger persistence and retention. |
tests/builtin/persistence/test_mysql_schema.py |
Budgets the new indexed datetime columns. |
tests/builtin/runtime/test_decision_composition.py |
Tests loopback endpoint classification. |
tests/builtin/runtime/test_decision_config.py |
Tests governance configuration defaults. |
tests/builtin/runtime/test_decision_policy.py |
Tests assessment invariants and mappings. |
tests/builtin/runtime/test_memory_write_gate_contract.py |
Tests policy modes, privacy, and observations. |
tests/builtin/runtime/test_memory_write_gate_paths.py |
Tests end-to-end gate and ledger behavior. |
tests/test_server_logging.py |
Tests serialization of decision log fields. |
🧠 Review effort: Balanced
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| operation_id=_gate_operation_id(base), | ||
| subject_refs=_gate_subject_refs(candidates), | ||
| evidence_refs=tuple(_gate_evidence_ref(entry) for entry in projection.entries), | ||
| observation_sink=self._write_gate_observation_sink, |
| self._decision_model = decision_model | ||
| self._hold_on = hold_on | ||
| self._threshold = threshold | ||
| self._policy = _MEMORY_WRITE_POLICY if policy is None else policy | ||
| self._allow_model_evaluation = ( |
| return tuple( | ||
| f"candidate:{index}" | ||
| if candidate.entry is None | ||
| else f"entry:{candidate.entry.entry_id}@{candidate.entry.entry_version_id}" | ||
| for index, candidate in enumerate(candidates, start=1) |
| _SAFE_REFERENCE = re.compile( | ||
| r"(?:candidate|evidence):[0-9]+|(?:candidate:[0-9]+ )?(?:entry|source|artifact):[A-Za-z0-9_.:@-]{1,500}\Z" | ||
| ) |
| if not enabled or mode is DecisionPolicyMode.DISABLED or decision_model is None: | ||
| return None |
| privacy_boundary=self._policy.privacy_boundary, | ||
| privacy_outcome=self._privacy_outcome(policy_assessment), | ||
| model_policy_id=None if policy_assessment.source.value == "none" else self.policy_id, | ||
| assessment=_safe_observation_assessment(policy_assessment), |


Which issue or RFC does this PR close? Implements the first bounded Memory write-gate slice of RFC #1770. Refs #1770, #1649, #1643, #1644, #1645, #1647, #1648, #1742, and #1745. This PR does not close the broader implementation or evaluation issues. # Rationale for this change RFC #1770 requires an explicit policy, privacy, fallback, observation, and promotion contract before additional domain consumers adopt decision models. Memory write evidence sufficiency is the narrowest existing boundary: the gate may produce a visible
HOLD, but Memory retains authority over commits and revisions. The design borrows only bounded governance principles from Hermes/JEV research: explicit modes, deterministic local rules, safe fallback, and observable outcomes. It does not copy Hermes storage, lifecycle, prompts, thresholds, or host-specific implementation. # What changes are included in this PR? - Adds runtime-independentDecisionPolicy,DecisionAssessment, andDecisionObservationcontracts. A fallback is alwaysunadjudicated. - Adds an opt-in Memory evidence-sufficiency gate withdisabled -> shadow -> advisory -> enforcingbehavior. Runtime configuration defaults toshadow; enforcing is an explicit operator choice. - Binds the gate's hold direction and confidence threshold into the effective immutable policy version, and declares labeled calibration, measured friction/cost/latency, and rollback as promotion criteria. - Applies mode semantics and a private observation to deterministic evidence-projection rejections as well as model decisions. Existing custom gates without the optional preflight contract retain conservative legacy behavior. - Enforces the privacy boundary on every gate construction path. Raw candidate/evidence text is evaluated only whenlocal_onlyis selected and the resolved backend is explicitly loopback (localhost,127.0.0.1, or::1). Unknown, LAN, wildcard, and hosted endpoints fail open without receiving raw content.no_external_call,references_only, and unsupportedhosted_redacteddo not make a content call. - Emits bounded structured logs and a private, scope-isolated durablepc_decision_observationsledger. Observations retain policy/version/mode, stable references, counts, verdict/action, privacy boundary/outcome, and fallback codes only; they never store candidate text, evidence text, model rationale, or exception details. - Adds exact observation replay idempotency, same-ID conflict rejection, and optionaldecision_observation_retention_dayscleanup limited to the current Scope. The default is no automatic deletion. - Completes MySQL schema/index budgeting and bilingualobloaderrecovery lists for the new table. # Are there any user-facing changes? No public HTTP/OpenAPI endpoint, CLI command, dashboard, export surface, hosted-redacted transport, replay/rescore job, query rewrite, or additional decision-model consumer is added. The Memory write gate remains disabled unless explicitly enabled. When enabled, its runtime default isshadowand its privacy default isno_external_call; it only records a no-call observation and preserves the write. An operator must explicitly selectlocal_only, configure a loopback backend, calibrate the policy, and chooseenforcingbefore raw content is evaluated and a write can be held. # How was this change tested? -uv run --no-sync pytest tests/builtin/runtime/test_memory_write_gate_contract.py tests/builtin/runtime/test_memory_write_gate_paths.py tests/builtin/runtime/test_decision_config.py tests/builtin/runtime/test_decision_composition.py tests/builtin/runtime/test_decision_policy.py tests/test_server_logging.py tests/builtin/persistence/test_mysql_schema.py -q(108 passed) -uv run --no-sync ruff check <16 changed files>-uv run --no-sync ruff format --check <16 changed files>-uv run --no-sync ty check --python-version 3.11 <14 changed Python files>-git diff --checkValidation limitation: the broadertests/builtin/runtime -qrun consistently stops in the unchanged Windows spawned-worker diagnostic case expectingworker_crashbut observingprocessing_failed(197 passed, 1 failedbefore--maxfail=1). This PR does not modify that worker or diagnostic path; the focused decision, persistence, logging, and composition suites above were run after the final changes. # AI usage statement Implemented with OpenAI Codex in the Codex desktop app from RFC #1770 and repository-local inspection. Hermes/JEV was consulted only as architectural prior art; no Hermes source or implementation structure was copied.Follow-up
Promotion to
enforcingis not part of this PR. The replay/re-evaluation, calibration, held-out evaluation, and promotion decision needed to consider it are tracked in #1930.