Skip to content

feat(runtime): add decision policy governance types - #1921

Open
AlexStocks wants to merge 6 commits into
masterfrom
codex/decision-policy-governance
Open

AlexStocks wants to merge 6 commits into
masterfrom
codex/decision-policy-governance

Conversation

@AlexStocks

@AlexStocks AlexStocks commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

Which issue or RFC does this PR close? Implements the first bounded Memory write-gate slice of RFC #1770. Refs #1770, #1649, #1643, #1644, #1645, #1647, #1648, #1742, and #1745. This PR does not close the broader implementation or evaluation issues. # Rationale for this change RFC #1770 requires an explicit policy, privacy, fallback, observation, and promotion contract before additional domain consumers adopt decision models. Memory write evidence sufficiency is the narrowest existing boundary: the gate may produce a visible HOLD, but Memory retains authority over commits and revisions. The design borrows only bounded governance principles from Hermes/JEV research: explicit modes, deterministic local rules, safe fallback, and observable outcomes. It does not copy Hermes storage, lifecycle, prompts, thresholds, or host-specific implementation. # What changes are included in this PR? - Adds runtime-independent DecisionPolicy, DecisionAssessment, and DecisionObservation contracts. A fallback is always unadjudicated. - Adds an opt-in Memory evidence-sufficiency gate with disabled -> shadow -> advisory -> enforcing behavior. Runtime configuration defaults to shadow; enforcing is an explicit operator choice. - Binds the gate's hold direction and confidence threshold into the effective immutable policy version, and declares labeled calibration, measured friction/cost/latency, and rollback as promotion criteria. - Applies mode semantics and a private observation to deterministic evidence-projection rejections as well as model decisions. Existing custom gates without the optional preflight contract retain conservative legacy behavior. - Enforces the privacy boundary on every gate construction path. Raw candidate/evidence text is evaluated only when local_only is selected and the resolved backend is explicitly loopback (localhost, 127.0.0.1, or ::1). Unknown, LAN, wildcard, and hosted endpoints fail open without receiving raw content. no_external_call, references_only, and unsupported hosted_redacted do not make a content call. - Emits bounded structured logs and a private, scope-isolated durable pc_decision_observations ledger. Observations retain policy/version/mode, stable references, counts, verdict/action, privacy boundary/outcome, and fallback codes only; they never store candidate text, evidence text, model rationale, or exception details. - Adds exact observation replay idempotency, same-ID conflict rejection, and optional decision_observation_retention_days cleanup limited to the current Scope. The default is no automatic deletion. - Completes MySQL schema/index budgeting and bilingual obloader recovery lists for the new table. # Are there any user-facing changes? No public HTTP/OpenAPI endpoint, CLI command, dashboard, export surface, hosted-redacted transport, replay/rescore job, query rewrite, or additional decision-model consumer is added. The Memory write gate remains disabled unless explicitly enabled. When enabled, its runtime default is shadow and its privacy default is no_external_call; it only records a no-call observation and preserves the write. An operator must explicitly select local_only, configure a loopback backend, calibrate the policy, and choose enforcing before raw content is evaluated and a write can be held. # How was this change tested? - uv run --no-sync pytest tests/builtin/runtime/test_memory_write_gate_contract.py tests/builtin/runtime/test_memory_write_gate_paths.py tests/builtin/runtime/test_decision_config.py tests/builtin/runtime/test_decision_composition.py tests/builtin/runtime/test_decision_policy.py tests/test_server_logging.py tests/builtin/persistence/test_mysql_schema.py -q (108 passed) - uv run --no-sync ruff check <16 changed files> - uv run --no-sync ruff format --check <16 changed files> - uv run --no-sync ty check --python-version 3.11 <14 changed Python files> - git diff --check Validation limitation: the broader tests/builtin/runtime -q run consistently stops in the unchanged Windows spawned-worker diagnostic case expecting worker_crash but observing processing_failed (197 passed, 1 failed before --maxfail=1). This PR does not modify that worker or diagnostic path; the focused decision, persistence, logging, and composition suites above were run after the final changes. # AI usage statement Implemented with OpenAI Codex in the Codex desktop app from RFC #1770 and repository-local inspection. Hermes/JEV was consulted only as architectural prior art; no Hermes source or implementation structure was copied.

Follow-up

Promotion to enforcing is not part of this PR. The replay/re-evaluation, calibration, held-out evaluation, and promotion decision needed to consider it are tracked in #1930.

Add the internal DecisionPolicy, DecisionAssessment, and DecisionObservation vocabulary from RFC 1770. Route the Memory write gate through the shared assessment mapping before preserving its existing ACCEPT/FLAG/HOLD contract.

Constraint: no public HTTP/OpenAPI surface, persistence schema, or hosted-provider configuration is added.

Rejected: copying Hermes implementation code or plugin structure; only the RFC-level governance semantics are mirrored.

Scope-risk: limited to runtime decision vocabulary and Memory write gate assessment mapping.

Tested: UV_PROJECT_ENVIRONMENT=.venv312 uv run --python 3.12 --no-sync pytest tests/builtin/runtime/test_decision_policy.py tests/builtin/runtime/test_memory_write_gate_contract.py::test_memory_write_decision_policy_keeps_fallback_unadjudicated tests/builtin/runtime/test_memory_write_gate_contract.py::test_memory_write_decision_policy_uses_configured_hold_direction -q

Tested: UV_PROJECT_ENVIRONMENT=.venv312 uv run --python 3.12 --no-sync ruff check src/powercontext/builtin/runtime/decision_policy.py src/powercontext/builtin/runtime/memory_write_gate.py src/powercontext/builtin/runtime/__init__.py tests/builtin/runtime/test_decision_policy.py tests/builtin/runtime/test_memory_write_gate_contract.py

Tested: UV_PROJECT_ENVIRONMENT=.venv312 uv run --python 3.12 --no-sync ruff format --check src/powercontext/builtin/runtime/decision_policy.py src/powercontext/builtin/runtime/memory_write_gate.py src/powercontext/builtin/runtime/__init__.py tests/builtin/runtime/test_decision_policy.py tests/builtin/runtime/test_memory_write_gate_contract.py

Not-tested: full Memory write gate async contract on this Windows host hangs in CPython Proactor event-loop socketpair setup; default uv Python 3.14 currently fails tests/conftest.py typing import; ty check cannot resolve pydantic/pytest in this environment.

Co-authored-by: OmX <omx@oh-my-codex.dev>
Apply the RFC 1770 policy mode, privacy-boundary, and observation rules to the existing Memory write gate. Default configuration now prevents content egress, while local-only evaluation keeps the established opt-in enforcement path.

Constraint: no HTTP/OpenAPI contract, persistence schema, durable observation ledger, hosted sanitizer, or provider integration is added.

Scope-risk: limited to Memory write-gate configuration, decisions, privacy-safe observation logs, and request attribution.

Tested: uv run --no-sync pytest tests/builtin/runtime/test_memory_write_gate_contract.py -q (26 passed)

Tested: uv run --no-sync pytest tests/builtin/runtime/test_memory_write_gate_contract.py tests/builtin/runtime/test_memory_write_gate_paths.py -q (46 passed)

Tested: uv run --no-sync ruff check <11 changed files>; uv run --no-sync ruff format --check <11 changed files>; uv run --no-sync ty check --python-version 3.11 <11 changed files>

Not-tested: a final Windows full-runtime retry was stopped after exceeding the prior 175s run time. The only earlier full-suite failure was reproduced against the exact b1d544b baseline in the spawned-worker diagnostics path, unrelated to these files.

@hidb4ai hidb4ai left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed ba7c086. Three issues need addressing before merge: shadow/advisory can suppress Memory creation, the declared privacy boundary does not constrain the actual backend call, and the Server logger drops the new observation fields. Reproductions and required behavior are in the inline comments.

Comment thread src/powercontext/builtin/runtime/memory_write_gate.py
Comment thread src/powercontext/builtin/runtime/memory_write_gate.py Outdated
Comment thread src/powercontext/builtin/runtime/decision_policy.py
Persist a Scope-isolated, reference-only observation ledger for Memory write-gate decisions. Sidecar failures do not alter gate outcomes.

Constraint: observations exclude candidate content, evidence bodies, model rationale, and exception messages.

Scope-risk: retention prunes expired observations only within the Scope receiving a write.

Tested: uv run --no-sync pytest tests/builtin/persistence/test_decision_observation_repository.py tests/builtin/runtime/test_decision_config.py tests/builtin/runtime/test_memory_write_gate_contract.py tests/builtin/runtime/test_memory_write_gate_paths.py; uv run --no-sync ruff check <16 changed files>; uv run --no-sync ruff format --check <16 changed files>; uv run --no-sync ty check --python-version 3.11 <16 changed files>.
Apply mode semantics to evidence preflight rejections, enforce loopback-only local evaluation, and make privacy outcomes observable in the private decision ledger and JSON logs.

Constraint: observations retain bounded references and counts only; no candidate or evidence content is logged or stored.

Scope-risk: unknown, LAN, wildcard, and hosted backends fail open under local_only until an explicitly loopback endpoint is configured.

Tested: focused gate, policy, composition, logging, and MySQL schema suite (108 passed); Ruff, format, ty 3.11, and diff check.
@hidb4ai

hidb4ai commented Oct 10, 2026

Copy link
Copy Markdown

resolve conflicts

@hidb4ai hidb4ai left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-reviewed 62b4117 and replayed the original counterexamples: the shadow/advisory, privacy-boundary, and serialized-log findings are fixed. Two remaining issues in the new observation ledger need addressing; details and reproductions are inline.

Comment on lines +147 to +148
if not all(_SAFE_REFERENCE.fullmatch(value) for value in (*observation.subject_refs, *observation.evidence_refs)):
raise ValueError("decision observation references must be structured identifiers") # noqa: TRY003

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve observations for supported Source identifiers

The reference allowlist is narrower than the supported Source ID contract. Through public capture/remember APIs with real SQLite, https://example.test/page, src/settings.py, and 会议记录 all write Memory successfully but leave zero ledger observations: this validation raises ValueError, which the sink caller catches and logs. The control ID safe persists one observation. This makes audit coverage depend on identifier spelling and loses normal URL, path, and Unicode references. Preserve supported references in a bounded structured representation or safe encoding instead of discarding the entire observation.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Still reproduced on d4a8603 through public capture/remember APIs with real SQLite: safe writes Memory and one ledger observation, while https://example.test/page, src/settings.py, and 会议记录 each write Memory but persist zero observations. All three log decision.observation_sink_failed with ValueError. The current allowlist still rejects these supported Source IDs; the RFC and upstream-merge updates do not change this path. Please preserve these references in a bounded structured representation or safe encoding so the observation is retained.

privacy_outcome=self._privacy_outcome(policy_assessment),
model_policy_id=None if policy_assessment.source.value == "none" else self.policy_id,
assessment=_safe_observation_assessment(policy_assessment),
final_action=f"memory_write_{assessment.verdict.value}",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Record the domain outcome separately from gate acceptance

This final_action is persisted before Memory commit preparation/application finishes. With memory_max_active_entries=1, a second remember() raises MemoryCapacityExceededError and leaves revision 1, yet the ledger records memory-write:memory@2 with final_action=memory_write_accept. An unapplied plan and a deduplicated no-op produce the same action. ACCEPT correctly describes gate permission, but it does not supply RFC 1770's required final domain action actually taken. Keep that assessment and record the actual outcome at the Memory owner boundary so failed, unapplied, unchanged, and applied operations remain distinguishable.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Still reproduced on d4a8603: with memory_max_active_entries=1, the second remember() raises MemoryCapacityExceededError and Memory remains at revision 1, but the ledger records memory-write:memory@2 with final_action=memory_write_accept. Unapplied plans and deduplicated no-ops also record acceptance. The RFC update defers replay tooling but still requires the final domain action actually taken. This field is emitted before commit preparation/application and is never finalized; retain gate permission separately and record the actual outcome at the Memory owner boundary.

Clarify that replay and rescore are promotion evidence for enforcing policies, rather than mandatory runtime capabilities for every consumer. Track the Memory write-gate evidence work in #1930.

Constraint: the RFC does not require additional consumers, a dashboard, replay tooling, or a cross-consumer evaluation bundle.

Scope-risk: consumers without promotion evidence remain disabled, shadow, or advisory; existing authority, Scope, revision, and citation boundaries are unchanged.

Tested: website pnpm lint; website pnpm test (13 passed); website pnpm build (255 public pages and 405 repository files link-validated; 950 public pages export-verified); git diff --check.
…cy-governance

# Conflicts:
#	docs/en/docs/operate/troubleshoot.md
#	docs/zh/docs/operate/troubleshoot.md
@AlexStocks
AlexStocks marked this pull request as ready for review October 11, 2026 03:14
Copilot AI balanced review requested due to automatic review settings October 11, 2026 03:14

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Observation timing, reference encoding, policy consistency, and missing backend metadata currently violate the intended governance contract.

6 open findings
What changed in this PR

Adds governed decision-policy contracts and applies them to the Memory write gate with privacy modes, structured observations, and durable persistence.

Changes:

  • Introduces decision policy, assessment, privacy, and observation types.
  • Adds mode-aware Memory gating and scoped observation retention.
  • Updates schema, recovery docs, logging, and focused tests.
File Description
.env.example Documents gate modes and privacy settings.
docs/​en/​docs/​operate/​troubleshoot.md Adds observation-table recovery guidance.
docs/​en/​rfcs/​1770-decision-policy-governance.md Clarifies promotion and replay requirements.
docs/​zh/​docs/​operate/​troubleshoot.md Updates Chinese recovery guidance.
docs/​zh/​rfcs/​1770-decision-policy-governance.md Synchronizes RFC clarifications.
src/​powercontext/​builtin/​artifacts/​memory/​protocols.py Extends gate requests and preflight protocols.
src/​powercontext/​builtin/​artifacts/​memory/​service.py Supplies scoped references and observation sinks.
src/​powercontext/​builtin/​decision_observations.py Defines runtime-independent governance contracts.
src/​powercontext/​builtin/​persistence/​__init__.py Exports the observation repository.
src/​powercontext/​builtin/​persistence/​decision_observations.py Persists and validates observations.
src/​powercontext/​builtin/​persistence/​tables.py Defines the observation ledger schema.
src/​powercontext/​builtin/​runtime/​__init__.py Exports decision governance APIs.
src/​powercontext/​builtin/​runtime/​composition.py Composes privacy-aware decision backends.
src/​powercontext/​builtin/​runtime/​config.py Adds mode, privacy, and retention settings.
src/​powercontext/​builtin/​runtime/​decision_model.py Propagates backend-locality attestations.
src/​powercontext/​builtin/​runtime/​decision_observation_ledger.py Implements the relational observation sink.
src/​powercontext/​builtin/​runtime/​decision_policy.py Maps results into governed assessments.
src/​powercontext/​builtin/​runtime/​memory_write_gate.py Applies policy modes, privacy, and observations.
src/​powercontext/​builtin/​runtime/​relational.py Wires scoped ledger persistence into Memory.
src/​powercontext/​server/​logging.py Adds decision fields to structured logs.
tests/​builtin/​persistence/​test_decision_observation_repository.py Tests ledger persistence and retention.
tests/​builtin/​persistence/​test_mysql_schema.py Budgets the new indexed datetime columns.
tests/​builtin/​runtime/​test_decision_composition.py Tests loopback endpoint classification.
tests/​builtin/​runtime/​test_decision_config.py Tests governance configuration defaults.
tests/​builtin/​runtime/​test_decision_policy.py Tests assessment invariants and mappings.
tests/​builtin/​runtime/​test_memory_write_gate_contract.py Tests policy modes, privacy, and observations.
tests/​builtin/​runtime/​test_memory_write_gate_paths.py Tests end-to-end gate and ledger behavior.
tests/​test_server_logging.py Tests serialization of decision log fields.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +1389 to +1392
operation_id=_gate_operation_id(base),
subject_refs=_gate_subject_refs(candidates),
evidence_refs=tuple(_gate_evidence_ref(entry) for entry in projection.entries),
observation_sink=self._write_gate_observation_sink,
Comment on lines 175 to +179
self._decision_model = decision_model
self._hold_on = hold_on
self._threshold = threshold
self._policy = _MEMORY_WRITE_POLICY if policy is None else policy
self._allow_model_evaluation = (
Comment on lines +1861 to +1865
return tuple(
f"candidate:{index}"
if candidate.entry is None
else f"entry:{candidate.entry.entry_id}@{candidate.entry.entry_version_id}"
for index, candidate in enumerate(candidates, start=1)
Comment on lines +35 to +37
_SAFE_REFERENCE = re.compile(
r"(?:candidate|evidence):[0-9]+|(?:candidate:[0-9]+ )?(?:entry|source|artifact):[A-Za-z0-9_.:@-]{1,500}\Z"
)
Comment on lines +138 to 139
if not enabled or mode is DecisionPolicyMode.DISABLED or decision_model is None:
return None
Comment on lines +307 to +310
privacy_boundary=self._policy.privacy_boundary,
privacy_outcome=self._privacy_outcome(policy_assessment),
model_policy_id=None if policy_assessment.source.value == "none" else self.policy_id,
assessment=_safe_observation_assessment(policy_assessment),

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants