Skip to content
Open
Show file tree
Hide file tree
Changes from 4 commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .gitattributes
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
benchmark/locomo/dataset/locomo10.json text eol=lf
evaluation/memory/locomo/dataset/locomo10.json text eol=lf
e2e/bub/harbor-tasks/** text eol=lf
e2e/bub/tasks/*.yaml text eol=lf
integrations/dsh/plugins/powercontext/src/operations.generated.ts text eol=lf
Expand Down
4 changes: 3 additions & 1 deletion .github/CODEOWNERS
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,9 @@
/tests/ @Teingi @PsiACE
/e2e/ @PsiACE @AlexStocks
/evaluation/ @PsiACE @frf12
/benchmark/ @Teingi @Alanxtl
/evaluation/memory/ @Teingi @Alanxtl
/evaluation/memory/longmemeval_v2/ @PsiACE @frf12
/evaluation/performance/ @Teingi @Alanxtl

# Documentation, website, and examples.
/docs/ @zhanghuidinah @PsiACE
Expand Down
9 changes: 2 additions & 7 deletions .github/workflows/master.yml
Original file line number Diff line number Diff line change
Expand Up @@ -77,13 +77,8 @@ jobs:
- name: Set up the environment
uses: ./.github/actions/setup-python-env

# `evaluation/` is its own uv project with its own `testpaths`, so the root
# unit-test job never collected it and the benchmark's own tests were only
# ever run by hand. The job starts at the unit subtree because that is the
# subtree this change owns; `evaluation/tests/web` and
# `evaluation/tests/contract` are already red on `master` for reasons
# unrelated to the benchmark, and gating them here would make this check
# red before it could report anything.
# Evaluation has its own uv project. Collect each suite's unit tests through
# its dedicated path; the root project's testpaths do not include them.
- name: Run the evaluation project unit tests
run: |
set -o pipefail
Expand Down
8 changes: 4 additions & 4 deletions .github/workflows/skill-guidance.yml
Original file line number Diff line number Diff line change
Expand Up @@ -3,14 +3,14 @@ name: Skill guidance validation
on:
pull_request:
paths:
- evaluation/skill-up/**
- evaluation/skills/skill-up/**
- openapi/powercontext.yaml
- integrations/claude-code/plugins/powercontext/skills/powercontext-project-context/**
- .github/workflows/skill-guidance.yml
push:
branches: [master]
paths:
- evaluation/skill-up/**
- evaluation/skills/skill-up/**
- openapi/powercontext.yaml
- integrations/claude-code/plugins/powercontext/skills/powercontext-project-context/**
- .github/workflows/skill-guidance.yml
Expand All @@ -31,7 +31,7 @@ jobs:
with:
python-version: '3.11'
- name: Install validation dependencies
run: python -m pip install -r evaluation/skill-up/requirements-dev.txt
run: python -m pip install -r evaluation/skills/skill-up/requirements-dev.txt
- name: Install checksum-pinned skill-up
shell: bash
run: |
Expand All @@ -43,7 +43,7 @@ jobs:
tar -xzf skill-up.tar.gz
echo "$RUNNER_TEMP/skill-up" >> "$GITHUB_PATH"
- name: Verify pin, fixtures, reports and native schema
working-directory: evaluation/skill-up
working-directory: evaluation/skills/skill-up
run: |
python sync_skill.py --check
python validate_suite.py
Expand Down
8 changes: 4 additions & 4 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -27,10 +27,10 @@ build/
develop-eggs/
dist/
node_modules/
evaluation/web/node_modules/
evaluation/web/playwright-report/
evaluation/web/test-results/
evaluation/web/*.tsbuildinfo
evaluation/coding/swebench_pro/web/node_modules/
evaluation/coding/swebench_pro/web/playwright-report/
evaluation/coding/swebench_pro/web/test-results/
evaluation/coding/swebench_pro/web/*.tsbuildinfo
# Local native-code experiment and validation reports.
/docs/*/rfcs/0000-native-git-code-understanding-validation.md
/docs/*/rfcs/0000-native-git-code-understanding-experiment.md
Expand Down
6 changes: 3 additions & 3 deletions .licenserc.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -50,9 +50,9 @@ header:
- '**/*.lock'
- '**/*-lock.yaml'
- '**/*.lockb'
- 'benchmark/locomo/dataset/**'
- 'benchmark/locomo/results/**'
- 'benchmark/locomo_plus/results/**'
- 'evaluation/memory/locomo/dataset/**'
- 'evaluation/memory/locomo/results/**'
- 'evaluation/memory/locomo_plus/results/**'
- 'tox.ini'
- 'Makefile'
- 'zensical.toml'
Expand Down
4 changes: 2 additions & 2 deletions .pre-commit-config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -9,10 +9,10 @@ repos:
- id: check-json
exclude: ^.devcontainer/devcontainer.json
- id: pretty-format-json
exclude: ^(?:\.devcontainer/devcontainer\.json|benchmark/locomo/dataset/locomo10\.json)$
exclude: ^(?:\.devcontainer/devcontainer\.json|evaluation/memory/locomo/dataset/locomo10\.json)$
args: [--autofix, --no-sort-keys, --no-ensure-ascii]
- id: end-of-file-fixer
exclude: ^benchmark/locomo/dataset/locomo10\.json$
exclude: ^evaluation/memory/locomo/dataset/locomo10\.json$
- id: trailing-whitespace

- repo: https://github.com/astral-sh/ruff-pre-commit
Expand Down
5 changes: 4 additions & 1 deletion Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -49,7 +49,10 @@ e2e-test: ## Run CLI to Client SDK to Server end-to-end tests.
.PHONY: evaluation-unit-test
evaluation-unit-test: ## Run the evaluation project's unit tests, the work-continuity benchmark included.
@uv sync --project evaluation --frozen
@uv run --project evaluation pytest -c evaluation/pyproject.toml evaluation/tests/unit -m "not live" -q
@uv run --project evaluation pytest -c evaluation/pyproject.toml \
evaluation/coding/swebench_pro/tests/unit \
evaluation/coding/work_continuity/tests/unit \
evaluation/memory/longmemeval_v2/tests -m "not live" -q

.PHONY: code-seekdb-test
code-seekdb-test: ## Exercise native code indexing against a real embedded seekdb instance.
Expand Down
2 changes: 2 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -125,6 +125,8 @@ make test
```

See [CONTRIBUTING.md](CONTRIBUTING.md) for the complete development workflow.
Coding and memory evaluations, performance benchmarks, and Skill regressions live under
[`evaluation/`](evaluation/README.md).

## Learn more

Expand Down
1 change: 1 addition & 0 deletions README_CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -116,6 +116,7 @@ make test
```

完整开发流程请阅读 [CONTRIBUTING.md](CONTRIBUTING.md)。
编程能力评测、记忆质量评测、性能压测和 Skill 回归统一放在 [`evaluation/`](evaluation/README.md),按用途选择对应目录。

## 进一步了解

Expand Down
1 change: 1 addition & 0 deletions README_JP.md
Original file line number Diff line number Diff line change
Expand Up @@ -102,6 +102,7 @@ make test
```

開発ワークフロー全体については [CONTRIBUTING.md](CONTRIBUTING.md) を参照してください。
コーディング評価、メモリ品質評価、性能ベンチマーク、Skill 回帰テストは [`evaluation/`](evaluation/README.md) にまとめています。

## さらに詳しく

Expand Down
10 changes: 0 additions & 10 deletions benchmark/README.md

This file was deleted.

2 changes: 1 addition & 1 deletion docs/en/development/integration-guidance-evaluation.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ description: Measured routing, execution evidence, and remaining model limitatio

# Agent guidance evaluation record

The complementary [skill-up regression suite](https://github.com/oceanbase/powercontext/blob/master/evaluation/skill-up/README.md) pins the packaged Claude Code
The complementary [skill-up regression suite](https://github.com/oceanbase/powercontext/blob/master/evaluation/skills/skill-up/README.md) pins the packaged Claude Code
Skill and checks tool selection and authorization boundaries with rule-based assertions, controlled MCP replies, and
a with/without-Skill comparison. It covers Claude Code + MCP only, with hooks disabled and permissions bypassed;
it does not qualify real host approval, bounded recall, automatic Capture/Flush, persistence or memory quality.
Expand Down
2 changes: 1 addition & 1 deletion docs/en/development/releasing.md
Original file line number Diff line number Diff line change
Expand Up @@ -95,7 +95,7 @@ PowerContext release number.
| WorkBuddy integration | Its hooks and transport scripts carry User-Agent strings; there is no plugin manifest version to align with the Server. |
| Skill Receiver | `RECEIVER_VERSION` in `src/powercontext/client/skill_receiver.py` supplies the default receiver identity version reported during enrollment, reconciliation, and receipts. It is independent of the Server version. |
| Python integrations | `integrations/{bub,langchain,langgraph,opendal,pydantic-ai}/pyproject.toml` contain separate distribution versions. Dependency constraints express compatibility, not the current main release. |
| Evaluation and harness packages | `evaluation/pyproject.toml`, `evaluation/web/package.json`, and `e2e/bub/pyproject.toml` have their own package versions. |
| Evaluation and harness packages | `evaluation/pyproject.toml`, `evaluation/coding/swebench_pro/web/package.json`, and `e2e/bub/pyproject.toml` have their own package versions. |
| Protocols and dependencies | Agent Plugins schema versions, persisted format versions, OpenAPI specification version, API paths, host minimum versions, and third-party dependencies change only with their own contracts. |

Installing an Agent integration from the matching `powercontext-v…` repository tag selects the matching source
Expand Down
8 changes: 4 additions & 4 deletions docs/en/rfcs/1718_memory_capacity_contract.md
Original file line number Diff line number Diff line change
Expand Up @@ -238,7 +238,7 @@ the byte ceiling even at the default identifier widths.

The 5,000 / 10,000 / 4 MiB defaults remain configurable growth limits. Calibrate deployment budgets with isolated,
representative measurements on the selected backend, including retained-history cost. Benchmark methodology and
measurement summaries belong in `benchmark/memory_capacity/README.md`; raw run results accompany acceptance evidence.
measurement summaries belong in `evaluation/performance/memory_capacity/README.md`; raw run results accompany acceptance evidence.

## Configuration

Expand Down Expand Up @@ -501,8 +501,8 @@ projection work, so any change in their statement counts means the implementatio

### Scale benchmark

A new `benchmark/memory_capacity/` module, alongside the existing `locomo` benchmarks and outside `tests/` for the
reasons `benchmark/README.md` gives, records at entry counts 200, 1,000, and 5,000, and across a compaction cycle:
The `evaluation/performance/memory_capacity/` module sits outside `tests/` for the reasons
`evaluation/README.md` gives. It records at entry counts 200, 1,000, and 5,000, and across a compaction cycle:

entry count, manifest bytes, database bytes, mean append latency, mean final-window append latency, projection row
writes per append, and search recall behavior before and after compaction.
Expand All @@ -524,7 +524,7 @@ Each step is independently reviewable and leaves the tree green.
addition, `make api-generate`, `make contract-test`.
5. **Read bounding.** The `revisions()` cap and its capability error.
6. **Public read endpoint.** `POST /v1/memory/capacity`, OpenAPI, contract test.
7. **Benchmark.** `benchmark/memory_capacity/` and the recorded SQLite and OceanBase results.
7. **Benchmark.** `evaluation/performance/memory_capacity/` and the recorded SQLite and OceanBase results.

Steps 1 through 3 alone close the "no observable ceiling" half of #1718 and are worth landing before compaction.

Expand Down
2 changes: 1 addition & 1 deletion docs/zh/development/integration-guidance-evaluation.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ description: tracking issue 1450 D 的工具选择测量、执行证据及模型

# Agent 指引验证记录

配套的 [skill-up 回归套件](https://github.com/oceanbase/powercontext/blob/master/evaluation/skill-up/README.md) 固定 Claude Code 打包 Skill 的版本,
配套的 [skill-up 回归套件](https://github.com/oceanbase/powercontext/blob/master/evaluation/skills/skill-up/README.md) 固定 Claude Code 打包 Skill 的版本,
使用规则断言、受控 MCP 响应及加载/不加载 Skill 的对比,检查工具选择与授权边界。
该套件仅覆盖 Claude Code + MCP;运行器关闭 hooks 并绕过权限提示,因此不证明真实宿主审批、
有界召回、自动 Capture/Flush、持久化或记忆质量。声明实测覆盖时,应同时保留真实模型的原始记录与输入快照。
Expand Down
2 changes: 1 addition & 1 deletion docs/zh/development/releasing.md
Original file line number Diff line number Diff line change
Expand Up @@ -91,7 +91,7 @@ OpenAPI 版本描述待发布的 API 契约,不覆盖 Hatch VCS。打 tag 前
| WorkBuddy 集成 | Hook 和传输脚本包含 User-Agent 字符串,没有需要与 Server 对齐的插件 manifest 版本。 |
| Skill Receiver | `src/powercontext/client/skill_receiver.py` 中的 `RECEIVER_VERSION` 提供注册、协调及回执上报时默认使用的 Receiver 身份版本,独立于 Server 版本。 |
| Python 集成包 | `integrations/{bub,langchain,langgraph,opendal,pydantic-ai}/pyproject.toml` 有各自的发行版本;依赖范围表示兼容性,不是当前主发布版本。 |
| 评测及 harness 包 | `evaluation/pyproject.toml`、`evaluation/web/package.json`、`e2e/bub/pyproject.toml` 有各自的包版本。 |
| 评测及 harness 包 | `evaluation/pyproject.toml`、`evaluation/coding/swebench_pro/web/package.json`、`e2e/bub/pyproject.toml` 有各自的包版本。 |
| 协议及依赖 | Agent Plugins schema 版本、持久化格式版本、OpenAPI 规范版本、API 路径、宿主最低版本及第三方依赖仅随自身契约更新。 |

通过匹配的 `powercontext-v…` 仓库 tag 安装 Agent 集成,可以选中配套的源码修订,不要求每个插件 manifest 的
Expand Down
8 changes: 4 additions & 4 deletions docs/zh/rfcs/1718_memory_capacity_contract.md
Original file line number Diff line number Diff line change
Expand Up @@ -217,7 +217,7 @@ class MemoryCapacity(BaseModel):
标识符宽度,大批量写入也可能先触发字节上限。

保留可配置的 5,000 / 10,000 / 4 MiB 增长上限。部署预算应通过对应后端上隔离、具有代表性的测量校准,并计入保留
历史的成本。基准方法与测量摘要集中在 `benchmark/memory_capacity/README.md`,原始运行结果随验收证据保存。
历史的成本。基准方法与测量摘要集中在 `evaluation/performance/memory_capacity/README.md`,原始运行结果随验收证据保存。

## 配置

Expand Down Expand Up @@ -457,8 +457,8 @@ PR #1709 的 `test_memory_append_projection_writes_do_not_grow_with_entry_histor

### 规模基准

新增 `benchmark/memory_capacity/` 模块,与现有 `locomo` 基准并列、并按 `benchmark/README.md` 给出的理由置于
`tests/` 之外,在 entry 数 200、1,000、5,000 以及一个完整压缩周期上记录:
`evaluation/performance/memory_capacity/` 模块按 `evaluation/README.md` 给出的理由置于 `tests/` 之外,
在 entry 数 200、1,000、5,000 以及一个完整压缩周期上记录:

entry 数、manifest 字节数、数据库字节数、平均 append 延迟、末窗平均 append 延迟、每次 append 的投影行写入数,
以及压缩前后的搜索召回行为。
Expand All @@ -479,7 +479,7 @@ entry 数、manifest 字节数、数据库字节数、平均 append 延迟、末
`make contract-test`。
5. **读取约束。** `revisions()` 的上限及其 capability 错误。
6. **公开读取端点。** `POST /v1/memory/capacity`、OpenAPI、契约测试。
7. **基准。** `benchmark/memory_capacity/` 以及记录的 SQLite 与 OceanBase 结果。
7. **基准。** `evaluation/performance/memory_capacity/` 以及记录的 SQLite 与 OceanBase 结果。

仅第 1 至 3 步就能闭合 #1718 中"没有可观测上限"的那一半,值得在压缩之前先合入。

Expand Down
2 changes: 1 addition & 1 deletion e2e/bub/tasks/locomo-multihop-football.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ categories:
- sample
- batch:locomo
provenance:
source: benchmark/locomo/dataset/locomo10.json
source: evaluation/memory/locomo/dataset/locomo10.json
revision: 4448275ea2c5cd0af5774d80aea7b05b5a16e1b996caf8554ca3d762a301ae84
selection: category-3-first-two-explicit-named-facts/v1
case_ids:
Expand Down
2 changes: 1 addition & 1 deletion e2e/bub/tasks/locomo-open-pastries.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ categories:
- sample
- batch:locomo
provenance:
source: benchmark/locomo/dataset/locomo10.json
source: evaluation/memory/locomo/dataset/locomo10.json
revision: 4448275ea2c5cd0af5774d80aea7b05b5a16e1b996caf8554ca3d762a301ae84
selection: category-4-first-three-item-list/v1
case_ids:
Expand Down
2 changes: 1 addition & 1 deletion e2e/bub/tasks/locomo-support-group.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ categories:
- sample
- batch:locomo
provenance:
source: benchmark/locomo/dataset/locomo10.json
source: evaluation/memory/locomo/dataset/locomo10.json
revision: 4448275ea2c5cd0af5774d80aea7b05b5a16e1b996caf8554ca3d762a301ae84
selection: first-conversation-first-question/v1
case_ids:
Expand Down
2 changes: 1 addition & 1 deletion e2e/bub/tasks/locomo-temporal-banker.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ categories:
- sample
- batch:locomo
provenance:
source: benchmark/locomo/dataset/locomo10.json
source: evaluation/memory/locomo/dataset/locomo10.json
revision: 4448275ea2c5cd0af5774d80aea7b05b5a16e1b996caf8554ca3d762a301ae84
selection: category-2-first-short-answer-after-conv-26/v1
case_ids:
Expand Down
2 changes: 1 addition & 1 deletion e2e/bub/tests/test_harbor_job_config.py
Original file line number Diff line number Diff line change
Expand Up @@ -520,7 +520,7 @@ def test_agent_container_cannot_read_workload_answers(monkeypatch, tmp_path: Pat
monkeypatch.setenv(plugin_host.server_url, "http://host-gateway:8000")
task = load_tasks(_REPOSITORY / "e2e" / "bub" / manifest)[0]
protected = [_REPOSITORY / "e2e" / "bub" / name for name in ("harbor-tasks", "paired-tasks", "tasks")]
protected.append(_REPOSITORY / "benchmark")
protected.append(_REPOSITORY / "evaluation")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Heads with #1853: e2e/bub/tests/test_harbor_job_config.py conflicts, and one side is a security-relevant path

This file is the only conflict between the two branches, but it is worth resolving deliberately rather than by taking one side. git merge-tree 673d44d6 <this head> 799ea12c reports exactly one changed in both:

  base   100644 b9fe8124...  e2e/bub/tests/test_harbor_job_config.py
  our    100644 293e50a6...  e2e/bub/tests/test_harbor_job_config.py
  their  100644 06a40cbe...  e2e/bub/tests/test_harbor_job_config.py

This PR changes one line there, and it is the one that matters:

-    protected.append(_REPOSITORY / "benchmark")
+    protected.append(_REPOSITORY / "evaluation")

That is the protected list handed to the agent container in test_agent_container_cannot_read_workload_answers. Since this PR moves benchmark/ to evaluation/, dropping that line would leave the answers directory readable by the agent under test — the assertion would still be testing something, but not the protection.

#1853 rewrites the same file (about 120 added lines in test_pi_runs_without_a_saved_session, plus imports and five other tests), so a textual conflict is unavoidable; the two changes do not conflict semantically. Whichever side lands first, please keep both: the protected entry pointing at evaluation, and #1853's new tests.

Also worth re-checking after the merge: #1853 has an open review comment on this file (4192361330) whose line anchors were computed against its own head.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in eb5b658 by merging upstream master. All of #1853's tests are preserved; the only difference from the upstream test file is the protected path changing from benchmark to evaluation. I also verified that the mount-configuration checks reject an injected evaluation/ mount and that the existing-proxy regression passes with uppercase and lowercase HTTP/HTTPS proxy variables set.


sources = [Path(mount["source"]) for mount in _config(task, tmp_path, host=host_adapter(host)).environment.mounts]

Expand Down
3 changes: 3 additions & 0 deletions e2e/bub/uv.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

Loading
Loading