Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .gitattributes
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
benchmark/locomo/dataset/locomo10.json text eol=lf
evaluation/memory/locomo/dataset/locomo10.json text eol=lf
e2e/bub/harbor-tasks/** text eol=lf
e2e/bub/tasks/*.yaml text eol=lf
integrations/dsh/plugins/powercontext/src/operations.generated.ts text eol=lf
Expand Down
4 changes: 3 additions & 1 deletion .github/CODEOWNERS
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,9 @@
/tests/ @Teingi @PsiACE
/e2e/ @PsiACE @AlexStocks
/evaluation/ @PsiACE @frf12
/benchmark/ @Teingi @Alanxtl
/evaluation/memory/ @Teingi @Alanxtl
/evaluation/memory/longmemeval_v2/ @PsiACE @frf12
/evaluation/performance/ @Teingi @Alanxtl

# Documentation, website, and examples.
/docs/ @zhanghuidinah @PsiACE
Expand Down
9 changes: 2 additions & 7 deletions .github/workflows/master.yml
Original file line number Diff line number Diff line change
Expand Up @@ -77,13 +77,8 @@ jobs:
- name: Set up the environment
uses: ./.github/actions/setup-python-env

# `evaluation/` is its own uv project with its own `testpaths`, so the root
# unit-test job never collected it and the benchmark's own tests were only
# ever run by hand. The job starts at the unit subtree because that is the
# subtree this change owns; `evaluation/tests/web` and
# `evaluation/tests/contract` are already red on `master` for reasons
# unrelated to the benchmark, and gating them here would make this check
# red before it could report anything.
# Evaluation has its own uv project. Collect each suite's unit tests through
# its dedicated path; the root project's testpaths do not include them.
- name: Run the evaluation project unit tests
run: |
set -o pipefail
Expand Down
8 changes: 4 additions & 4 deletions .github/workflows/skill-guidance.yml
Original file line number Diff line number Diff line change
Expand Up @@ -3,14 +3,14 @@ name: Skill guidance validation
on:
pull_request:
paths:
- evaluation/skill-up/**
- evaluation/skills/skill-up/**
- openapi/powercontext.yaml
- integrations/claude-code/plugins/powercontext/skills/powercontext-project-context/**
- .github/workflows/skill-guidance.yml
push:
branches: [master]
paths:
- evaluation/skill-up/**
- evaluation/skills/skill-up/**
- openapi/powercontext.yaml
- integrations/claude-code/plugins/powercontext/skills/powercontext-project-context/**
- .github/workflows/skill-guidance.yml
Expand All @@ -31,7 +31,7 @@ jobs:
with:
python-version: '3.11'
- name: Install validation dependencies
run: python -m pip install -r evaluation/skill-up/requirements-dev.txt
run: python -m pip install -r evaluation/skills/skill-up/requirements-dev.txt
- name: Install checksum-pinned skill-up
shell: bash
run: |
Expand All @@ -43,7 +43,7 @@ jobs:
tar -xzf skill-up.tar.gz
echo "$RUNNER_TEMP/skill-up" >> "$GITHUB_PATH"
- name: Verify pin, fixtures, reports and native schema
working-directory: evaluation/skill-up
working-directory: evaluation/skills/skill-up
run: |
python sync_skill.py --check
python validate_suite.py
Expand Down
8 changes: 4 additions & 4 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -27,10 +27,10 @@ build/
develop-eggs/
dist/
node_modules/
evaluation/web/node_modules/
evaluation/web/playwright-report/
evaluation/web/test-results/
evaluation/web/*.tsbuildinfo
evaluation/coding/swebench_pro/web/node_modules/
evaluation/coding/swebench_pro/web/playwright-report/
evaluation/coding/swebench_pro/web/test-results/
evaluation/coding/swebench_pro/web/*.tsbuildinfo
# Local native-code experiment and validation reports.
/docs/*/rfcs/0000-native-git-code-understanding-validation.md
/docs/*/rfcs/0000-native-git-code-understanding-experiment.md
Expand Down
6 changes: 3 additions & 3 deletions .licenserc.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -50,9 +50,9 @@ header:
- '**/*.lock'
- '**/*-lock.yaml'
- '**/*.lockb'
- 'benchmark/locomo/dataset/**'
- 'benchmark/locomo/results/**'
- 'benchmark/locomo_plus/results/**'
- 'evaluation/memory/locomo/dataset/**'
- 'evaluation/memory/locomo/results/**'
- 'evaluation/memory/locomo_plus/results/**'
- 'tox.ini'
- 'Makefile'
- 'zensical.toml'
Expand Down
4 changes: 2 additions & 2 deletions .pre-commit-config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -9,10 +9,10 @@ repos:
- id: check-json
exclude: ^.devcontainer/devcontainer.json
- id: pretty-format-json
exclude: ^(?:\.devcontainer/devcontainer\.json|benchmark/locomo/dataset/locomo10\.json)$
exclude: ^(?:\.devcontainer/devcontainer\.json|evaluation/memory/locomo/dataset/locomo10\.json)$
args: [--autofix, --no-sort-keys, --no-ensure-ascii]
- id: end-of-file-fixer
exclude: ^benchmark/locomo/dataset/locomo10\.json$
exclude: ^evaluation/memory/locomo/dataset/locomo10\.json$
- id: trailing-whitespace

- repo: https://github.com/astral-sh/ruff-pre-commit
Expand Down
5 changes: 4 additions & 1 deletion Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -49,7 +49,10 @@ e2e-test: ## Run CLI to Client SDK to Server end-to-end tests.
.PHONY: evaluation-unit-test
evaluation-unit-test: ## Run the evaluation project's unit tests, the work-continuity benchmark included.
@uv sync --project evaluation --frozen
@uv run --project evaluation pytest -c evaluation/pyproject.toml evaluation/tests/unit -m "not live" -q
@uv run --project evaluation pytest -c evaluation/pyproject.toml \
evaluation/coding/swebench_pro/tests/unit \
evaluation/coding/work_continuity/tests/unit \
evaluation/memory/longmemeval_v2/tests -m "not live" -q

.PHONY: code-seekdb-test
code-seekdb-test: ## Exercise native code indexing against a real embedded seekdb instance.
Expand Down
2 changes: 2 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -125,6 +125,8 @@ make test
```

See [CONTRIBUTING.md](CONTRIBUTING.md) for the complete development workflow.
Coding and memory evaluations, performance benchmarks, and Skill regressions live under
[`evaluation/`](evaluation/README.md).

## Learn more

Expand Down
1 change: 1 addition & 0 deletions README_CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -116,6 +116,7 @@ make test
```

完整开发流程请阅读 [CONTRIBUTING.md](CONTRIBUTING.md)。
编程能力评测、记忆质量评测、性能压测和 Skill 回归统一放在 [`evaluation/`](evaluation/README.md),按用途选择对应目录。

## 进一步了解

Expand Down
1 change: 1 addition & 0 deletions README_JP.md
Original file line number Diff line number Diff line change
Expand Up @@ -102,6 +102,7 @@ make test
```

開発ワークフロー全体については [CONTRIBUTING.md](CONTRIBUTING.md) を参照してください。
コーディング評価、メモリ品質評価、性能ベンチマーク、Skill 回帰テストは [`evaluation/`](evaluation/README.md) にまとめています。

## さらに詳しく

Expand Down
10 changes: 0 additions & 10 deletions benchmark/README.md

This file was deleted.

2 changes: 1 addition & 1 deletion docs/en/development/integration-guidance-evaluation.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ description: Measured routing, execution evidence, and remaining model limitatio

# Agent guidance evaluation record

The complementary [skill-up regression suite](https://github.com/oceanbase/powercontext/blob/master/evaluation/skill-up/README.md) pins the packaged Claude Code
The complementary [skill-up regression suite](https://github.com/oceanbase/powercontext/blob/master/evaluation/skills/skill-up/README.md) pins the packaged Claude Code
Skill and checks tool selection and authorization boundaries with rule-based assertions, controlled MCP replies, and
a with/without-Skill comparison. It covers Claude Code + MCP only, with hooks disabled and permissions bypassed;
it does not qualify real host approval, bounded recall, automatic Capture/Flush, persistence or memory quality.
Expand Down
2 changes: 1 addition & 1 deletion docs/en/development/releasing.md
Original file line number Diff line number Diff line change
Expand Up @@ -95,7 +95,7 @@ PowerContext release number.
| WorkBuddy integration | Its hooks and transport scripts carry User-Agent strings; there is no plugin manifest version to align with the Server. |
| Skill Receiver | `RECEIVER_VERSION` in `src/powercontext/client/skill_receiver.py` supplies the default receiver identity version reported during enrollment, reconciliation, and receipts. It is independent of the Server version. |
| Python integrations | `integrations/{bub,langchain,langgraph,opendal,pydantic-ai}/pyproject.toml` contain separate distribution versions. Dependency constraints express compatibility, not the current main release. |
| Evaluation and harness packages | `evaluation/pyproject.toml`, `evaluation/web/package.json`, and `e2e/bub/pyproject.toml` have their own package versions. |
| Evaluation and harness packages | `evaluation/pyproject.toml`, `evaluation/coding/swebench_pro/web/package.json`, and `e2e/bub/pyproject.toml` have their own package versions. |
| Protocols and dependencies | Agent Plugins schema versions, persisted format versions, OpenAPI specification version, API paths, host minimum versions, and third-party dependencies change only with their own contracts. |

Installing an Agent integration from the matching `powercontext-v…` repository tag selects the matching source
Expand Down
8 changes: 4 additions & 4 deletions docs/en/rfcs/1718_memory_capacity_contract.md
Original file line number Diff line number Diff line change
Expand Up @@ -238,7 +238,7 @@ the byte ceiling even at the default identifier widths.

The 5,000 / 10,000 / 4 MiB defaults remain configurable growth limits. Calibrate deployment budgets with isolated,
representative measurements on the selected backend, including retained-history cost. Benchmark methodology and
measurement summaries belong in `benchmark/memory_capacity/README.md`; raw run results accompany acceptance evidence.
measurement summaries belong in `evaluation/performance/memory_capacity/README.md`; raw run results accompany acceptance evidence.

## Configuration

Expand Down Expand Up @@ -501,8 +501,8 @@ projection work, so any change in their statement counts means the implementatio

### Scale benchmark

A new `benchmark/memory_capacity/` module, alongside the existing `locomo` benchmarks and outside `tests/` for the
reasons `benchmark/README.md` gives, records at entry counts 200, 1,000, and 5,000, and across a compaction cycle:
The `evaluation/performance/memory_capacity/` module sits outside `tests/` for the reasons
`evaluation/README.md` gives. It records at entry counts 200, 1,000, and 5,000, and across a compaction cycle:

entry count, manifest bytes, database bytes, mean append latency, mean final-window append latency, projection row
writes per append, and search recall behavior before and after compaction.
Expand All @@ -524,7 +524,7 @@ Each step is independently reviewable and leaves the tree green.
addition, `make api-generate`, `make contract-test`.
5. **Read bounding.** The `revisions()` cap and its capability error.
6. **Public read endpoint.** `POST /v1/memory/capacity`, OpenAPI, contract test.
7. **Benchmark.** `benchmark/memory_capacity/` and the recorded SQLite and OceanBase results.
7. **Benchmark.** `evaluation/performance/memory_capacity/` and the recorded SQLite and OceanBase results.

Steps 1 through 3 alone close the "no observable ceiling" half of #1718 and are worth landing before compaction.

Expand Down
2 changes: 1 addition & 1 deletion docs/zh/development/integration-guidance-evaluation.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ description: tracking issue 1450 D 的工具选择测量、执行证据及模型

# Agent 指引验证记录

配套的 [skill-up 回归套件](https://github.com/oceanbase/powercontext/blob/master/evaluation/skill-up/README.md) 固定 Claude Code 打包 Skill 的版本,
配套的 [skill-up 回归套件](https://github.com/oceanbase/powercontext/blob/master/evaluation/skills/skill-up/README.md) 固定 Claude Code 打包 Skill 的版本,
使用规则断言、受控 MCP 响应及加载/不加载 Skill 的对比,检查工具选择与授权边界。
该套件仅覆盖 Claude Code + MCP;运行器关闭 hooks 并绕过权限提示,因此不证明真实宿主审批、
有界召回、自动 Capture/Flush、持久化或记忆质量。声明实测覆盖时,应同时保留真实模型的原始记录与输入快照。
Expand Down
2 changes: 1 addition & 1 deletion docs/zh/development/releasing.md
Original file line number Diff line number Diff line change
Expand Up @@ -91,7 +91,7 @@ OpenAPI 版本描述待发布的 API 契约,不覆盖 Hatch VCS。打 tag 前
| WorkBuddy 集成 | Hook 和传输脚本包含 User-Agent 字符串,没有需要与 Server 对齐的插件 manifest 版本。 |
| Skill Receiver | `src/powercontext/client/skill_receiver.py` 中的 `RECEIVER_VERSION` 提供注册、协调及回执上报时默认使用的 Receiver 身份版本,独立于 Server 版本。 |
| Python 集成包 | `integrations/{bub,langchain,langgraph,opendal,pydantic-ai}/pyproject.toml` 有各自的发行版本;依赖范围表示兼容性,不是当前主发布版本。 |
| 评测及 harness 包 | `evaluation/pyproject.toml`、`evaluation/web/package.json`、`e2e/bub/pyproject.toml` 有各自的包版本。 |
| 评测及 harness 包 | `evaluation/pyproject.toml`、`evaluation/coding/swebench_pro/web/package.json`、`e2e/bub/pyproject.toml` 有各自的包版本。 |
| 协议及依赖 | Agent Plugins schema 版本、持久化格式版本、OpenAPI 规范版本、API 路径、宿主最低版本及第三方依赖仅随自身契约更新。 |

通过匹配的 `powercontext-v…` 仓库 tag 安装 Agent 集成,可以选中配套的源码修订,不要求每个插件 manifest 的
Expand Down
8 changes: 4 additions & 4 deletions docs/zh/rfcs/1718_memory_capacity_contract.md
Original file line number Diff line number Diff line change
Expand Up @@ -217,7 +217,7 @@ class MemoryCapacity(BaseModel):
标识符宽度,大批量写入也可能先触发字节上限。

保留可配置的 5,000 / 10,000 / 4 MiB 增长上限。部署预算应通过对应后端上隔离、具有代表性的测量校准,并计入保留
历史的成本。基准方法与测量摘要集中在 `benchmark/memory_capacity/README.md`,原始运行结果随验收证据保存。
历史的成本。基准方法与测量摘要集中在 `evaluation/performance/memory_capacity/README.md`,原始运行结果随验收证据保存。

## 配置

Expand Down Expand Up @@ -457,8 +457,8 @@ PR #1709 的 `test_memory_append_projection_writes_do_not_grow_with_entry_histor

### 规模基准

新增 `benchmark/memory_capacity/` 模块,与现有 `locomo` 基准并列、并按 `benchmark/README.md` 给出的理由置于
`tests/` 之外,在 entry 数 200、1,000、5,000 以及一个完整压缩周期上记录:
`evaluation/performance/memory_capacity/` 模块按 `evaluation/README.md` 给出的理由置于 `tests/` 之外,
在 entry 数 200、1,000、5,000 以及一个完整压缩周期上记录:

entry 数、manifest 字节数、数据库字节数、平均 append 延迟、末窗平均 append 延迟、每次 append 的投影行写入数,
以及压缩前后的搜索召回行为。
Expand All @@ -479,7 +479,7 @@ entry 数、manifest 字节数、数据库字节数、平均 append 延迟、末
`make contract-test`。
5. **读取约束。** `revisions()` 的上限及其 capability 错误。
6. **公开读取端点。** `POST /v1/memory/capacity`、OpenAPI、契约测试。
7. **基准。** `benchmark/memory_capacity/` 以及记录的 SQLite 与 OceanBase 结果。
7. **基准。** `evaluation/performance/memory_capacity/` 以及记录的 SQLite 与 OceanBase 结果。

仅第 1 至 3 步就能闭合 #1718 中"没有可观测上限"的那一半,值得在压缩之前先合入。

Expand Down
2 changes: 1 addition & 1 deletion e2e/bub/tasks/locomo-multihop-football.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ categories:
- sample
- batch:locomo
provenance:
source: benchmark/locomo/dataset/locomo10.json
source: evaluation/memory/locomo/dataset/locomo10.json
revision: 4448275ea2c5cd0af5774d80aea7b05b5a16e1b996caf8554ca3d762a301ae84
selection: category-3-first-two-explicit-named-facts/v1
case_ids:
Expand Down
2 changes: 1 addition & 1 deletion e2e/bub/tasks/locomo-open-pastries.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ categories:
- sample
- batch:locomo
provenance:
source: benchmark/locomo/dataset/locomo10.json
source: evaluation/memory/locomo/dataset/locomo10.json
revision: 4448275ea2c5cd0af5774d80aea7b05b5a16e1b996caf8554ca3d762a301ae84
selection: category-4-first-three-item-list/v1
case_ids:
Expand Down
2 changes: 1 addition & 1 deletion e2e/bub/tasks/locomo-support-group.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ categories:
- sample
- batch:locomo
provenance:
source: benchmark/locomo/dataset/locomo10.json
source: evaluation/memory/locomo/dataset/locomo10.json
revision: 4448275ea2c5cd0af5774d80aea7b05b5a16e1b996caf8554ca3d762a301ae84
selection: first-conversation-first-question/v1
case_ids:
Expand Down
2 changes: 1 addition & 1 deletion e2e/bub/tasks/locomo-temporal-banker.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ categories:
- sample
- batch:locomo
provenance:
source: benchmark/locomo/dataset/locomo10.json
source: evaluation/memory/locomo/dataset/locomo10.json
revision: 4448275ea2c5cd0af5774d80aea7b05b5a16e1b996caf8554ca3d762a301ae84
selection: category-2-first-short-answer-after-conv-26/v1
case_ids:
Expand Down
2 changes: 1 addition & 1 deletion e2e/bub/tests/test_harbor_job_config.py
Original file line number Diff line number Diff line change
Expand Up @@ -693,7 +693,7 @@ def test_agent_container_cannot_read_workload_answers(monkeypatch, tmp_path: Pat
monkeypatch.setenv(plugin_host.server_url, "http://host-gateway:8000")
task = load_tasks(_REPOSITORY / "e2e" / "bub" / manifest)[0]
protected = [_REPOSITORY / "e2e" / "bub" / name for name in ("harbor-tasks", "paired-tasks", "tasks")]
protected.append(_REPOSITORY / "benchmark")
protected.append(_REPOSITORY / "evaluation")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Heads with #1853: e2e/bub/tests/test_harbor_job_config.py conflicts, and one side is a security-relevant path

This file is the only conflict between the two branches, but it is worth resolving deliberately rather than by taking one side. git merge-tree 673d44d6 <this head> 799ea12c reports exactly one changed in both:

  base   100644 b9fe8124...  e2e/bub/tests/test_harbor_job_config.py
  our    100644 293e50a6...  e2e/bub/tests/test_harbor_job_config.py
  their  100644 06a40cbe...  e2e/bub/tests/test_harbor_job_config.py

This PR changes one line there, and it is the one that matters:

-    protected.append(_REPOSITORY / "benchmark")
+    protected.append(_REPOSITORY / "evaluation")

That is the protected list handed to the agent container in test_agent_container_cannot_read_workload_answers. Since this PR moves benchmark/ to evaluation/, dropping that line would leave the answers directory readable by the agent under test — the assertion would still be testing something, but not the protection.

#1853 rewrites the same file (about 120 added lines in test_pi_runs_without_a_saved_session, plus imports and five other tests), so a textual conflict is unavoidable; the two changes do not conflict semantically. Whichever side lands first, please keep both: the protected entry pointing at evaluation, and #1853's new tests.

Also worth re-checking after the merge: #1853 has an open review comment on this file (4192361330) whose line anchors were computed against its own head.


sources = [Path(mount["source"]) for mount in _config(task, tmp_path, host=host_adapter(host)).environment.mounts]

Expand Down
3 changes: 3 additions & 0 deletions e2e/bub/uv.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

Loading
Loading