Skip to content
Open
Show file tree
Hide file tree
Changes from 6 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -109,6 +109,14 @@ POWERCONTEXT_SERVER_RUNTIME_DREAM_MAX_PENDING_PER_SCOPE=32
# Memory write gate: opt-in evidence-sufficiency check before a write commits. Disabled by
# default; when unset no gate runs and no extra model call is made.
# POWERCONTEXT_SERVER_RUNTIME_MEMORY_WRITE_GATE_ENABLED=true
# `disabled` does not construct a gate. `shadow` records an internal observation but always lets
# the Memory write proceed. `advisory` turns an insufficient-evidence result into FLAG, while
# `enforcing` preserves the visible ACCEPT/FLAG/HOLD outcomes.
# POWERCONTEXT_SERVER_RUNTIME_MEMORY_WRITE_GATE_MODE=shadow
# The default makes no decision-model call. Set `local_only` only for an in-process or loopback
# decision backend controlled by this deployment. `hosted_redacted` is rejected until a shared
# PowerContext content sanitizer is separately reviewed.
# POWERCONTEXT_SERVER_RUNTIME_MEMORY_WRITE_GATE_PRIVACY_BOUNDARY=local_only

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] The example value contradicts the actual default for the privacy boundary

RuntimeConfig.memory_write_gate_privacy_boundary defaults to no_external_call (runtime/config.py, and test_memory_write_gate_defaults_to_no_external_call pins it), and the comment above this line correctly says "The default makes no decision-model call". But by this file's own convention the commented line shows the default (..._MODE=shadow two lines up matches its actual default), and here it shows local_only - a value that DOES make a decision-model call on a loopback backend.

An operator who uncomments the line to "keep the documented default" actually enables model calls. Suggest changing the example to # POWERCONTEXT_SERVER_RUNTIME_MEMORY_WRITE_GATE_PRIVACY_BOUNDARY=no_external_call (and optionally a follow-up line showing the local_only opt-in), or rewording so the example is clearly an opt-in and not the default.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 5098f5d. .env.example now shows no_external_call, matching RuntimeConfig and the surrounding documentation. local_only remains described as the explicit loopback opt-in.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 5098f5d. .env.example now shows no_external_call, matching RuntimeConfig and the surrounding documentation. local_only remains described as the explicit loopback opt-in.

# Direction only (which verdict means the cited evidence is insufficient): "yes" or "no".
# POWERCONTEXT_SERVER_RUNTIME_MEMORY_WRITE_GATE_HOLD_ON=yes
# Optional confidence floor below which a hold becomes a written-but-annotated change.
Expand Down
6 changes: 3 additions & 3 deletions docs/en/docs/operate/troubleshoot.md
Original file line number Diff line number Diff line change
Expand Up @@ -259,15 +259,15 @@ so the previous database remains available for recovery:
Because `pc_scopes.parent_scope_id` is self-referential, keep ancestor Scope rows before their descendants in the
exported `pc_scopes` data.
If the source predates the three Skill lifecycle tables (`pc_skill_packages`, `pc_agent_skill_targets`, and
`pc_skill_publications`), the Profile tables, `pc_topic_memory_work_budgets`, or
`pc_receipt_migration_review`, remove the absent tables from their respective layers.
`pc_skill_publications`), the Profile tables, `pc_topic_memory_work_budgets`, `pc_recall_effort_daily`,
`pc_receipt_migration_review`, or `pc_decision_observations`, remove the absent tables from their respective layers.
When a work-budget table exists, restore it together with Cursors so failure allowances survive the migration.

Layer 1 contains parents and tables without foreign keys:

```bash
obloader <connection-options> -D <new-database> --csv \
--table 'pc_scopes,pc_source_journal_heads,pc_sources,pc_artifacts,pc_source_cursors,pc_memory_source_windows,pc_artifact_processing_leases,pc_artifact_processing_binding_states,pc_artifact_processing_pending,pc_artifact_processing_auto_wave_targets,pc_artifact_processing_sequences,pc_artifact_processing_intents,pc_topic_memory_processing_targets,pc_artifact_processing_schema,pc_artifact_processing_migration_receipts,pc_topic_memory_work_budgets,pc_topic_memory_retrieval_shape,pc_connector_checkpoints,pc_source_definition_manifests,pc_external_skill_registrations,pc_skill_packages,pc_agent_skill_targets,pc_skill_publications,pc_model_usage_daily,pc_recall_token_daily,pc_recall_effort_daily,pc_receipt_migration_review' \
--table 'pc_scopes,pc_source_journal_heads,pc_sources,pc_artifacts,pc_source_cursors,pc_memory_source_windows,pc_artifact_processing_leases,pc_artifact_processing_binding_states,pc_artifact_processing_pending,pc_artifact_processing_auto_wave_targets,pc_artifact_processing_sequences,pc_artifact_processing_intents,pc_topic_memory_processing_targets,pc_artifact_processing_schema,pc_artifact_processing_migration_receipts,pc_topic_memory_work_budgets,pc_topic_memory_retrieval_shape,pc_connector_checkpoints,pc_source_definition_manifests,pc_external_skill_registrations,pc_skill_packages,pc_agent_skill_targets,pc_skill_publications,pc_model_usage_daily,pc_recall_token_daily,pc_recall_effort_daily,pc_receipt_migration_review,pc_decision_observations' \
-f <export-directory>
```

Expand Down
18 changes: 10 additions & 8 deletions docs/en/rfcs/1770-decision-policy-governance.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,15 +58,17 @@ behaviour. RFC 1652 reserves Memory consolidation, supersession, and lifecycle e
shared policy and observation contract, each consumer will reinvent its own threshold, fallback, privacy, and audit
rules.

The expected outcome is a small shared layer that lets implementation PRs proceed in a consistent order:
For a policy that seeks promotion beyond `shadow` or `advisory`, the shared layer supports this evidence sequence:

1. define a policy in shadow mode;
2. collect observations without changing domain behaviour;
3. replay and rescore the observations under candidate thresholds;
4. promote only the narrow action a policy has earned;
5. keep domain authority, review, persistence, and access control in the owning service.

This is an enablement RFC, not an adoption recommendation for a specific backend.
This is an enablement RFC, not an adoption recommendation for a specific backend or a commitment to implement every
listed consumer. A consumer may stop at `disabled`, `shadow`, or `advisory`; it does not need a replay runner, dashboard,
or cross-consumer evaluation bundle unless it seeks a stronger mode.

# Guide-level explanation

Expand Down Expand Up @@ -350,8 +352,8 @@ Shadow mode requirements:
2. Every gate attempt emits a `DecisionObservation`.
3. Observations distinguish no-call, local-rule, successful backend, backend fallback, privacy fallback, and parse
failure.
4. The policy can be replayed or rescored without re-running the whole workload when enough exact references remain
resolvable.
4. A policy considered for `enforcing` can be replayed or rescored without re-running the whole workload when enough
exact references remain resolvable.

Replay has two forms:

Expand All @@ -361,6 +363,10 @@ Replay has two forms:
Both forms must report their input coverage. A replay result that cannot resolve enough original material is not a
negative result for the policy; it is an incomplete replay.

Replay/rescore is required promotion evidence, not a required runtime feature for every `shadow` or `advisory` consumer.
A consumer without that evidence must not be promoted to `enforcing`.
Memory write-gate promotion evidence is tracked separately in [#1930](https://github.com/oceanbase/powercontext/issues/1930).

## Promotion criteria

A policy must declare promotion criteria before moving beyond shadow.
Expand Down Expand Up @@ -484,9 +490,6 @@ are candidate backends or reference points, not policy owners.

# Unresolved questions

- What is the first concrete storage target for `DecisionObservation`: structured logs, evaluation artifacts, or a
private internal ledger?
- Which subset of observations should be retained by default, and for how long?
- Should policy definitions live as Python constants, JSON/YAML resources, or both?
- Which numeric promotion criteria should #1742 use if it is revised to conform to this RFC?
- Do hosted-provider privacy boundaries need a user-facing configuration document before the first hosted backend is
Expand All @@ -499,6 +502,5 @@ are candidate backends or reference points, not policy owners.
- A policy registry page in the dashboard showing active policies, modes, fallback rates, and last replay results.
- A CLI command that replays one policy against recent observations and prints candidate thresholds.
- A standard evaluation bundle for decision policies, shared by Jev, Laya, local rules, and future backends.
- A private decision-observation ledger with retention controls and export for offline evaluation.
- Policy-aware Review Inbox grouping, where a review item links back to the policy observation that suggested it.
- Local-only decision backends for air-gapped deployments, using the same policy and observation contract.
9 changes: 4 additions & 5 deletions docs/zh/docs/operate/troubleshoot.md
Original file line number Diff line number Diff line change
Expand Up @@ -250,16 +250,15 @@ collation,但不会包含数据库 URL 或凭据。
删除它。由于 `pc_scopes.parent_scope_id` 自引用 `pc_scopes`,导出的 `pc_scopes` 数据必须让祖先 Scope 记录排在
后代记录之前。
源数据库早于三张 Skill 生命周期表(`pc_skill_packages`、`pc_agent_skill_targets` 和
`pc_skill_publications`)、Profile 表或 `pc_receipt_migration_review` 时,应从对应层删除缺失的表。

旧版本没有 `pc_topic_memory_work_budgets` 时可从清单移除该表;若该表存在,必须与 Cursor 一起恢复,
保留已消耗的失败额度,避免迁移后重置累计成本。
`pc_skill_publications`)、Profile 表、`pc_topic_memory_work_budgets`、`pc_recall_effort_daily`、
`pc_receipt_migration_review` 或 `pc_decision_observations` 时,应从对应层删除缺失的表。
如果存在 work-budget 表,必须与 Cursor 一起恢复,保留已消耗的失败额度,避免迁移后重置累计成本。

第 1 层包含父表和无外键的表:

```bash
obloader <connection-options> -D <new-database> --csv \
--table 'pc_scopes,pc_source_journal_heads,pc_sources,pc_artifacts,pc_source_cursors,pc_memory_source_windows,pc_artifact_processing_leases,pc_artifact_processing_binding_states,pc_artifact_processing_pending,pc_artifact_processing_auto_wave_targets,pc_artifact_processing_sequences,pc_artifact_processing_intents,pc_topic_memory_processing_targets,pc_artifact_processing_schema,pc_artifact_processing_migration_receipts,pc_topic_memory_work_budgets,pc_topic_memory_retrieval_shape,pc_connector_checkpoints,pc_source_definition_manifests,pc_external_skill_registrations,pc_skill_packages,pc_agent_skill_targets,pc_skill_publications,pc_model_usage_daily,pc_recall_token_daily,pc_recall_effort_daily,pc_receipt_migration_review' \
--table 'pc_scopes,pc_source_journal_heads,pc_sources,pc_artifacts,pc_source_cursors,pc_memory_source_windows,pc_artifact_processing_leases,pc_artifact_processing_binding_states,pc_artifact_processing_pending,pc_artifact_processing_auto_wave_targets,pc_artifact_processing_sequences,pc_artifact_processing_intents,pc_topic_memory_processing_targets,pc_artifact_processing_schema,pc_artifact_processing_migration_receipts,pc_topic_memory_work_budgets,pc_topic_memory_retrieval_shape,pc_connector_checkpoints,pc_source_definition_manifests,pc_external_skill_registrations,pc_skill_packages,pc_agent_skill_targets,pc_skill_publications,pc_model_usage_daily,pc_recall_token_daily,pc_recall_effort_daily,pc_receipt_migration_review,pc_decision_observations' \
-f <export-directory>
```

Expand Down
15 changes: 9 additions & 6 deletions docs/zh/rfcs/1770-decision-policy-governance.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,15 +44,17 @@

项目已经在这个边界上有具体压力。#1742 提出了 Memory 写入时的 evidence gate。#1745 提出了 decision-model reranker,并明确把 rerank 的 fail-closed 行为与 advisory gate 的 fail-open 行为分开。RFC 1652 将 Memory consolidation、supersession 和 lifecycle effects 留给 review-gated 流程。如果没有共享 policy 与 observation 契约,每个消费者都会重新发明自己的阈值、fallback、隐私和审计规则。

预期结果是一层很小的共享机制,使实现 PR 能按一致顺序推进:
对于寻求从 `shadow` 或 `advisory` 晋升的 policy,这层共享机制支持以下证据链:

1. 先用 shadow mode 定义 policy;
2. 在不改变领域行为的前提下收集 observation;
3. 用历史 observation replay 和 rescore 候选阈值;
4. 只 promote 该 policy 已经证明可以承担的窄域动作;
5. 领域权威、review、持久化和访问控制仍由所属服务负责。

这是一个 enablement RFC,不是对某个具体后端的采用建议。
这是一个 enablement RFC,不是对某个具体后端的采用建议,也不承诺实现列出的每个 consumer。consumer 可以停在
`disabled`、`shadow` 或 `advisory`,除非要进入更强模式,否则不需要实现 replay runner、dashboard 或跨 consumer
evaluation bundle。

# Guide-level explanation

Expand Down Expand Up @@ -285,7 +287,7 @@ Shadow mode 要求:
1. 最终领域输出与 no-gate 路径 byte-for-byte 或语义等价。
2. 每次 gate attempt 都产生 `DecisionObservation`。
3. observation 区分 no-call、local-rule、successful backend、backend fallback、privacy fallback 和 parse failure。
4. 只要仍能解析足够精确引用,policy 就可以在不重跑整个 workload 的情况下 replay 或 rescore。
4. 被考虑进入 `enforcing` 的 policy,在仍能解析足够精确引用时,可以不重跑整个 workload 而 replay 或 rescore。

Replay 有两种形式:

Expand All @@ -294,6 +296,10 @@ Replay 有两种形式:

两者都必须报告输入覆盖率。无法解析足够原始材料的 replay 结果不是 policy 的负面结果;它是不完整 replay。

replay/rescore 是 promotion 所需的证据,不是每个 `shadow` 或 `advisory` consumer 都必须交付的运行时功能。缺少该
证据的 consumer 不得晋升到 `enforcing`。
Memory write gate 的 promotion evidence 另由 [#1930](https://github.com/oceanbase/powercontext/issues/1930) 跟踪。

## Promotion criteria

policy 在超越 shadow 前必须先声明 promotion criteria。
Expand Down Expand Up @@ -394,8 +400,6 @@ PowerContext 已经有相关基础:

# Unresolved questions

- `DecisionObservation` 的第一个具体存储目标是什么:结构化日志、evaluation artifact,还是私有内部 ledger?
- 默认应保留哪些 observation,保留多久?
- policy definition 应以 Python constant、JSON/YAML resource,还是两者并存?
- 如果 #1742 调整为符合本 RFC,它应使用哪些数字 promotion criteria?
- 第一个托管后端启用前,是否需要先有面向用户的 hosted-provider privacy configuration 文档?
Expand All @@ -406,6 +410,5 @@ PowerContext 已经有相关基础:
- 在 dashboard 中增加 policy registry 页面,展示 active policy、mode、fallback rate 和最近 replay 结果。
- 增加 CLI 命令,对最近 observation replay 一个 policy,并输出候选阈值。
- 为 decision policy 提供标准 evaluation bundle,供 Jev、Laya、本地规则和未来后端共用。
- 增加带 retention controls 的私有 decision-observation ledger,并支持导出做离线 evaluation。
- Review Inbox 支持按 policy 分组,让 review item 反链到触发它的 policy observation。
- 面向气隙部署的 local-only decision backend,复用相同 policy 与 observation 契约。
31 changes: 30 additions & 1 deletion src/powercontext/builtin/artifacts/memory/protocols.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@
from contextlib import AbstractAsyncContextManager
from dataclasses import dataclass
from enum import StrEnum
from typing import Protocol
from typing import TYPE_CHECKING, Protocol, runtime_checkable

from pydantic import BaseModel, ConfigDict

Expand All @@ -39,6 +39,18 @@
from powercontext.builtin.tags import TagFilter
from powercontext.sources import Source

if TYPE_CHECKING:
from powercontext.builtin.decision_observations import DecisionObservation


class MemoryWriteObservationSink(Protocol):
"""Accept a bounded decision sidecar without owning the Memory write."""

async def record(self, observation: DecisionObservation, /) -> None:
"""Store one decision observation after the gate has made its judgement."""

...


class MemoryCandidateRequest(BaseModel):
"""Canonical evidence and bounded current entries offered to a pipeline."""
Expand Down Expand Up @@ -110,6 +122,11 @@ class MemoryWriteGateRequest:
candidates: tuple[str, ...]
evidence: tuple[str, ...]
expected_revision: int | None = None
scope_id: str = "unscoped"
operation_id: str | None = None
subject_refs: tuple[str, ...] = ()
evidence_refs: tuple[str, ...] = ()
observation_sink: MemoryWriteObservationSink | None = None


class MemoryWriteGate(Protocol):
Expand All @@ -127,6 +144,18 @@ async def assess(self, request: MemoryWriteGateRequest, /) -> MemoryWriteAssessm
...


@runtime_checkable
class MemoryWriteGatePreflight(Protocol):
"""Optionally map a deterministic service-side preflight rejection through a gate policy."""

async def assess_preflight(
self, request: MemoryWriteGateRequest, rejection: MemoryWriteAssessment, /
) -> MemoryWriteAssessment:
"""Return the mode-aware decision for a bounded preflight rejection."""

...


class MemoryWritePlan(BaseModel):
"""A side-effect-free result that can be committed in an outer transaction."""

Expand Down
45 changes: 38 additions & 7 deletions src/powercontext/builtin/artifacts/memory/service.py
Original file line number Diff line number Diff line change
Expand Up @@ -92,7 +92,9 @@
MemorySearchRequest,
MemoryWriteAssessment,
MemoryWriteGate,
MemoryWriteGatePreflight,
MemoryWriteGateRequest,
MemoryWriteObservationSink,
MemoryWritePlan,
MemoryWriteRejectionCode,
MemoryWriteVerdict,
Expand Down Expand Up @@ -259,7 +261,9 @@ def __init__(
artifact_resolver: _ArtifactResolver | None = None,
id_factory: IdFactory | None = None,
prompt_context: ScopedPrompts | None = None,
scope_id: str = "unscoped",
write_gate: MemoryWriteGate | None = None,
write_gate_observation_sink: MemoryWriteObservationSink | None = None,
capacity_budget: MemoryCapacityBudget | None = None,
compaction: MemoryCompactionPolicy | None = None,
max_history_revisions: int = 100,
Expand All @@ -268,6 +272,8 @@ def __init__(
self._prompt_context = prompt_context
self._candidate_pipeline = candidate_pipeline
self._write_gate = write_gate
self._write_gate_observation_sink = write_gate_observation_sink
self._scope_id = scope_id
self._embedding_model = embedding_model
if rerank_candidate_limit < 1:
raise _InvalidMemoryOperationError("search-limit")
Expand Down Expand Up @@ -1375,17 +1381,23 @@ async def _assess_write(
if self._write_gate is None:
return None
projection = await self._gate_evidence(base, candidates, evidence, current_entries)
request = MemoryWriteGateRequest(
candidates=tuple(candidate.text for candidate in candidates),
evidence=projection.entries,
expected_revision=None if base is None else base.revision,
scope_id=self._scope_id,
operation_id=_gate_operation_id(base),
subject_refs=_gate_subject_refs(candidates),
evidence_refs=tuple(_gate_evidence_ref(entry) for entry in projection.entries),
observation_sink=self._write_gate_observation_sink,
Comment thread
Copilot marked this conversation as resolved.
Outdated
)
if projection.rejection is not None:
if isinstance(self._write_gate, MemoryWriteGatePreflight):
return await self._write_gate.assess_preflight(request, projection.rejection)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] assess_preflight is not covered by the fail-open fallback that protects assess

_assess_write wraps only the assess() call in try/except Exception -> ACCEPT (a few lines below), but the new assess_preflight delegation returns directly with no exception handling. A gate that implements the optional MemoryWriteGatePreflight contract and raises during preflight (a bug, or a transient error inside a custom implementation) propagates out of _assess_write and fails the whole plan_remember, while the same gate's assess path fails open by design (failure_policy=FAIL_OPEN, and the PR's stated behavior is that gate problems never block writes).

Before this PR the projection-rejection path never called gate code, so this is a new failure mode introduced by the preflight hook. Suggest wrapping the preflight call in the same fail-open fallback (return an ACCEPT assessment with used_fallback=True and log the error type), so both entry points of a preflight-capable gate share one failure policy. A focused test: a preflight gate whose assess_preflight raises should still yield a committable plan.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 5098f5d. assess_preflight now shares the fail-open behavior of assess(): any gate exception returns an ACCEPT assessment with used_fallback=True, leaving the plan committable. Added a focused raising-preflight regression test.

_log_gate_assessment(projection.rejection)
return projection.rejection
try:
return await self._write_gate.assess(
MemoryWriteGateRequest(
candidates=tuple(candidate.text for candidate in candidates),
evidence=projection.entries,
expected_revision=None if base is None else base.revision,
)
)
return await self._write_gate.assess(request)
except Exception:
return MemoryWriteAssessment(
verdict=MemoryWriteVerdict.ACCEPT,
Expand Down Expand Up @@ -1839,6 +1851,25 @@ def _candidate_gate_identity(candidate_index: int, identity: str) -> str:
return f"candidate:{candidate_index} {identity}"


def _gate_operation_id(base: Memory | None) -> str:
if base is None:
return "memory-write:new"
return f"memory-write:{base.artifact_id}@{base.revision + 1}"


def _gate_subject_refs(candidates: tuple[MemoryEntryInput, ...]) -> tuple[str, ...]:
return tuple(
f"candidate:{index}"
if candidate.entry is None
else f"entry:{candidate.entry.entry_id}@{candidate.entry.entry_version_id}"
for index, candidate in enumerate(candidates, start=1)
Comment on lines +1987 to +1991

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 5098f5d. Applied writes replace candidate placeholders with exact committed entry/version refs and the committed Memory revision operation ID. Held, unchanged, and failed candidates cannot be reconstructed without retaining raw input, so their sidecars explicitly clear subject_refs and set incomplete_subject_count rather than claiming replayability.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Retain references for the complete assessed candidate batch

Applied single-candidate writes now have exact committed refs, but batch coverage is still incomplete on cd2bc83. Through MemoryService.remember() with real SQLite, storing First fact. and then writing [First fact., Second fact.] evaluates two candidates but persists only the newly created Second fact. entry in subject_refs; metadata says candidate_count=2 and has no incomplete_subject_count.

apply() uses only plan.commit.entry_versions, which excludes deduplicated/unchanged candidates and loses their ordinal mapping to evidence. Preserve exact refs for those candidates too, or explicitly record incomplete subjects and their positions, so re-evaluation can reconstruct the input actually judged.

)


def _gate_evidence_ref(value: str) -> str:
return value.partition("\n")[0]


def _source_gate_content(source: Source, resolver: _SourceResolver | None) -> str | None:
content = getattr(source, "content", None)
if isinstance(content, str):
Expand Down
Loading
Loading