Repository navigation
test(zcode): automate CLI acceptance and skill behavior evaluation #1850
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
jackie-cqz
wants to merge
5
commits into
oceanbase:master
Choose a base branch
from
jackie-cqz:fix/zcode-ci-acceptance
base: master
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
+2,351
−24
Open
Changes from 4 commits
Commits
Show all changes
5 commits
Select commit
Hold shift + click to select a range
f2329e9
test(zcode): automate CLI acceptance and skill behavior evaluation
jackie-cqz d49d27f
fix(ci): run ZCode guidance tests through Python
jackie-cqz 4df02e3
test(zcode): smoke check live evaluation entry point
jackie-cqz a228b1c
fix(zcode): distinguish native rejections from incomplete evidence
jackie-cqz de18005
fix(zcode): gate late capture responses in acceptance
jackie-cqz File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,141 @@ | ||
| name: ZCode acceptance | ||
|
|
||
| on: | ||
| pull_request: | ||
| types: [opened, synchronize, reopened, ready_for_review] | ||
| push: | ||
| branches: [master] | ||
| workflow_dispatch: | ||
| schedule: | ||
| - cron: '23 3 * * 0' | ||
|
|
||
| permissions: | ||
| contents: read | ||
|
|
||
| concurrency: | ||
| group: zcode-acceptance-${{ github.event.pull_request.number || github.ref }} | ||
| cancel-in-progress: true | ||
|
|
||
| env: | ||
| # ZCode 3.14.3 contains CLI 0.16.9. Update this pin only with a new acceptance run. | ||
| ZCODE_REF: 29628c9acdb81b703bbd4080c207a0e7ce5e276e | ||
| ZCODE_CLI_BIN: ${{ github.workspace }}/.powercontext/zcode-ci-source/apps/zcode-cli/packages/cli/dist/zcode.cjs | ||
| CI: 'true' | ||
| HUSKY: '0' | ||
|
|
||
| jobs: | ||
| controlled-cli: | ||
| runs-on: windows-latest | ||
| timeout-minutes: 30 | ||
| defaults: | ||
| run: | ||
| shell: pwsh | ||
| steps: | ||
| - name: Check out PowerContext | ||
| uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7 | ||
| with: | ||
| persist-credentials: false | ||
|
|
||
| - name: Check out the supported ZCode host | ||
| uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7 | ||
| with: | ||
| repository: zai-org/ZCode | ||
| ref: ${{ env.ZCODE_REF }} | ||
| path: .powercontext/zcode-ci-source | ||
| persist-credentials: false | ||
|
|
||
| - name: Set up Python | ||
| uses: ./.github/actions/setup-python-env | ||
| with: | ||
| python-version: '3.12' | ||
|
|
||
| - name: Set up Node | ||
| uses: actions/setup-node@249970729cb0ef3589644e2896645e5dc5ba9c38 # v6 | ||
| with: | ||
| node-version: '24.15.0' | ||
|
|
||
| - name: Install the host package manager | ||
| run: npm install --global pnpm@10.33.2 | ||
|
|
||
| - name: Validate the ZCode Skill pin and behavior evaluation harness | ||
| run: | | ||
| uv run --locked --no-sync python -m evaluation.zcode_guidance.pin --check | ||
| if ($LASTEXITCODE -ne 0) { exit $LASTEXITCODE } | ||
| uv run --locked --no-sync python -m evaluation.zcode_guidance.run --help | ||
| if ($LASTEXITCODE -ne 0) { exit $LASTEXITCODE } | ||
| uv run --locked --no-sync python -m pytest evaluation/zcode_guidance/tests -q | ||
|
|
||
| - name: Record build provenance | ||
| run: | | ||
| $hostCommit = git -C .powercontext/zcode-ci-source rev-parse HEAD | ||
| if ($LASTEXITCODE -ne 0 -or $hostCommit -ne $env:ZCODE_REF) { throw 'Unexpected ZCode commit' } | ||
| $reportDir = Join-Path $env:RUNNER_TEMP 'zcode-ci-evidence' | ||
| New-Item -ItemType Directory -Path $reportDir | Out-Null | ||
| [ordered]@{ | ||
| repository_commit = (git rev-parse HEAD) | ||
| host_repository = 'zai-org/ZCode' | ||
| host_commit = $hostCommit | ||
| node_version = (node --version) | ||
| pnpm_version = (pnpm --version) | ||
| model_mode = 'fixture' | ||
| server_mode = 'real-sqlite' | ||
| } | ConvertTo-Json | Set-Content -LiteralPath (Join-Path $reportDir 'provenance.json') -Encoding utf8 | ||
|
|
||
| - name: Install frozen host build dependencies | ||
| run: >- | ||
| pnpm --dir .powercontext/zcode-ci-source | ||
| --filter zcode --filter zcode-cli --filter '@zcode/cli...' | ||
| install --frozen-lockfile --ignore-scripts | ||
|
|
||
| - name: Build the actual CLI | ||
| run: | | ||
| pnpm --dir .powercontext/zcode-ci-source exec turbo --skip-infer --cwd apps/zcode-cli run build --filter=@zcode/cli | ||
| if ($LASTEXITCODE -ne 0) { exit $LASTEXITCODE } | ||
| if (-not (Test-Path -LiteralPath $env:ZCODE_CLI_BIN -PathType Leaf)) { throw 'CLI bundle not built' } | ||
| node $env:ZCODE_CLI_BIN --version | ||
|
|
||
| - name: Run all plugin Node tests including native host discovery | ||
| run: node --test integrations/zcode/plugins/powercontext/tests/*.test.mjs | ||
|
|
||
| - name: Run two independent controlled CLI acceptances | ||
| run: >- | ||
| uv run --locked --no-sync pytest tests/e2e/test_zcode_host_acceptance.py | ||
| -m 'zcode_host_acceptance and not zcode_live_model' | ||
| --zcode-acceptance-output "$env:RUNNER_TEMP/zcode-acceptance" -q | ||
|
|
||
| - name: Retain only summaries and selected evidence | ||
| if: always() | ||
| run: | | ||
| $reportDir = Join-Path $env:RUNNER_TEMP 'zcode-ci-evidence' | ||
| $output = Join-Path $env:RUNNER_TEMP 'zcode-acceptance' | ||
| if (Test-Path -LiteralPath $output) { | ||
| foreach ($run in Get-ChildItem -LiteralPath $output -Directory) { | ||
| if ($run.Name -notmatch '^[0-9a-f]{32}$') { continue } | ||
| $summary = Join-Path $run.FullName 'summary.json' | ||
| if (-not (Test-Path -LiteralPath $summary -PathType Leaf)) { continue } | ||
| $target = Join-Path $reportDir $run.Name | ||
| New-Item -ItemType Directory -Force -Path $target | Out-Null | ||
| Copy-Item -LiteralPath $summary -Destination $target | ||
| $evidence = Join-Path $run.FullName 'evidence' | ||
| if (Test-Path -LiteralPath $evidence -PathType Container) { | ||
| $evidenceTarget = Join-Path $target 'evidence' | ||
| New-Item -ItemType Directory -Path $evidenceTarget | Out-Null | ||
| Get-ChildItem -LiteralPath $evidence -File -Filter '*.json' | | ||
| Copy-Item -Destination $evidenceTarget | ||
| } | ||
| $result = Get-Content -LiteralPath $summary -Raw | ConvertFrom-Json | ||
| "### Run $($run.Name): $($result.overall_status)" | Add-Content -LiteralPath $env:GITHUB_STEP_SUMMARY | ||
| foreach ($scenario in $result.scenarios) { | ||
| "- $($scenario.id): $($scenario.status)" | Add-Content -LiteralPath $env:GITHUB_STEP_SUMMARY | ||
| } | ||
| } | ||
| } | ||
|
|
||
| - name: Upload controlled acceptance evidence | ||
| if: always() | ||
| uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7 | ||
| with: | ||
| name: zcode-controlled-${{ github.run_id }}-${{ github.run_attempt }} | ||
| path: ${{ runner.temp }}/zcode-ci-evidence/ | ||
| if-no-files-found: warn | ||
| retention-days: 14 | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,125 @@ | ||
| # ZCode Skill behavior evaluation | ||
|
|
||
| This suite runs the packaged PowerContext Skill through the **actual open-source ZCode CLI app-server and a live | ||
| model**. It adapts the paired guidance evaluation used by [skill-up](../skill-up/README.md) to ZCode's native | ||
| session, plugin, permission and MCP protocols. It does not invoke skill-up's Claude Code harness. | ||
|
|
||
| Each case has a fresh workspace, user profile and controlled HTTP MCP fixture. Both arms receive the same complete | ||
| catalog projected by the real PowerContext Server from its current contract. Forbidden operations remain callable: | ||
| attempts are recorded and fail the relevant rule. The `without_skill` arm removes only the PowerContext Skill; MCP, | ||
| binding Hooks, model selection and prompts remain identical. Neither arm receives the real Server's additional MCP | ||
| guidance in `initialize`, keeping that guidance from obscuring the Skill comparison. | ||
|
|
||
| ## Cases and rules | ||
|
|
||
| | Case | Required behavior | | ||
| | --- | --- | | ||
| | `ordinary-coding` | Answer the self-contained Python task with `[1, 2]`; schedule no tools. | | ||
| | `explicit-save` | Positive control: call native `remember_memory` once in the resolved Scope, save the requested decision and report the confirmed success. | | ||
| | `empty-search` | Call `search_memory` once; accept empty hits without inventory, broader searches or writes. Answer a subsequent self-contained coding question without tools. | | ||
| | `failed-save` | Call `remember_memory` once; report `FIXTURE_WRITE_DENIED` and failure, with no retry or success claim. | | ||
| | `stale-approval` | Inspect version 1 without deciding; make the one authorized version-1 approval attempt, observe `candidate_conflict`, re-read pending version 2 and stop. Never retry approval under the old authorization. | | ||
|
|
||
| Scope resolution uses the installed read-only `scope.mjs` and the native Hook's current-request binding metadata. | ||
| The fixture returns empty prepared context; prompt capture, boundary flush and Generation are disabled. Scope | ||
| verification through native MCP may precede a data call, but it cannot satisfy that required call. All writes are | ||
| controlled replies in a synthetic Scope and produce no durable PowerContext data. | ||
|
|
||
| These cases measure **routing and authorization behavior**, separately from the | ||
| [real Server acceptance](../../integrations/zcode/acceptance/README.md). They establish neither memory quality nor | ||
| official Windows desktop behavior. Fixture MCP operations are pre-authorized within the disposable test; this suite | ||
| does not evaluate an interactive permission UI. | ||
|
|
||
| ## Validate offline | ||
|
|
||
| Run from the repository root with the normal development dependencies: | ||
|
|
||
| ```powershell | ||
| uv run --locked python -m evaluation.zcode_guidance.pin --check | ||
| uv run --locked python -m evaluation.zcode_guidance.run --help | ||
| uv run --locked python -m pytest evaluation/zcode_guidance/tests -q | ||
| ``` | ||
|
|
||
| The `run --help` smoke check loads the live runner and its shared acceptance imports, then validates its command-line | ||
| entry point without starting a host or calling a model. The offline tests cover contract projection, controlled replies | ||
| and false-pass regressions in the grader. They do not establish model behavior. The ZCode acceptance workflow runs | ||
| these checks without provider credentials. | ||
|
|
||
| `skill-lock.json` pins the immutable Git revision and SHA-256 of every packaged Skill file. A changed checkout fails | ||
| the pin check. To evaluate a deliberately updated, committed Skill: | ||
|
|
||
| ```powershell | ||
| uv run --locked python -m evaluation.zcode_guidance.pin --revision <committed-revision> | ||
| ``` | ||
|
|
||
| Update the pin together with new live evidence; changing the pin alone grants no qualification. | ||
|
|
||
| ## Run a live evaluation | ||
|
|
||
| Use Node.js and a built ZCode CLI. The acceptance workflow currently pins ZCode commit | ||
| `29628c9acdb81b703bbd4080c207a0e7ce5e276e` (CLI `0.16.9`, desktop `3.14.3`). A host upgrade needs a fresh run. | ||
| Configure a tool-capable model in ZCode first. Supply its native versioned `.zcode/v2/provider_config.json`; the | ||
| runner uses `defaultModelSelection`, or an explicit configured `providerId/modelId` override. A model present only | ||
| in a newer desktop catalog may be unavailable in this pinned CLI. | ||
|
|
||
| ```powershell | ||
| $env:ZCODE_CLI_BIN = 'C:\path\to\ZCode\apps\zcode-cli\packages\cli\dist\zcode.cjs' | ||
| uv run --locked python -m evaluation.zcode_guidance.run ` | ||
| --model-config "$env:USERPROFILE\.zcode\v2\provider_config.json" ` | ||
| --output '.powercontext/zcode-guidance/first-live-run' | ||
| ``` | ||
|
|
||
| For example, append `--host-model 'bigmodel-api/GLM-5.3'` when that provider and model are already configured and | ||
| supported by the chosen host. No API Key, Token or Authorization value belongs on the command line. | ||
|
|
||
| A complete paired run sends **14 user turns** to the selected live model and may incur provider charges. The output | ||
| directory must be new. Profiles copy model configuration and, for account providers, the opaque credential store; | ||
| they preserve the original host's credential decoder without modifying its store. Copies are removed after the | ||
| owned app-server exits. Keep the entire output private under ignored `.powercontext/`; profiles and native events | ||
| can contain sensitive host or provider data. The runner never uploads live evidence. | ||
|
|
||
| ## Evidence and replay | ||
|
|
||
| Retain `inputs/`, `provenance.json`, `manifest.json` and both arm directories: | ||
|
|
||
| - Inputs contain the exact prompts, pinned Skill bytes, complete MCP catalog and evaluation source snapshot. | ||
| - Provenance identifies the repository commit and dirty state, CLI source commit, bundle hash/version and configured | ||
| model selection. Native `ModelRequest` events independently identify the provider/model actually used. | ||
| - Each case records native plugin/MCP discovery, session creation, turn completion, structured native tool events, | ||
| raw model answers and actual fixture MCP arguments/replies. Failed initialization remains a failed run. | ||
| - The report compares scheduled native tools with received MCP calls. ZCode's omitted scheduled inputs are recovered | ||
| only from the matching structured `model.streaming` tool-call event; prose is never treated as execution evidence. | ||
|
|
||
| Regrade archived bytes without using today's prompts or Skill pin: | ||
|
|
||
| ```powershell | ||
| uv run --locked python -m evaluation.zcode_guidance.report ` | ||
| --run '.powercontext/zcode-guidance/first-live-run' ` | ||
| --output '.powercontext/zcode-guidance/rechecked-report.json' | ||
| ``` | ||
|
|
||
| SHA-256 checks detect changed or missing archived bytes; they are integrity checks, not signed attestations. Literal | ||
| answer rules are deliberately transparent, and raw responses remain necessary for semantic review. Equivalent plain, | ||
| inline-code and fenced-code list rendering is accepted. Regraded reports identify the actual grader's source hash. | ||
| CLI exit status | ||
| is zero only when both arms complete and all `with_skill` cases pass. Baseline results are reported separately; | ||
| baseline behavior failure does not fail the gate, while incomplete baseline execution does. A completed turn with | ||
| the pinned host's matching native input-schema rejection counts as a behavior failure; the rejected attempt still | ||
| counts toward routing and retry rules, but requires no MCP wire call. Missing rejection evidence, contradictory | ||
| handler events or unexplained native/wire differences leave execution incomplete. One paired run is not | ||
| statistical evidence that the Skill improves behavior. A passing run means the | ||
| listed cases passed for that exact host, model, Skill and contract, not that every possible instruction is safe. | ||
|
|
||
| CI runs the offline gate. Live model evaluation is an explicit maintainer run after changing the Skill, host pin or | ||
| public contract; retain its report before claiming live qualification. CI does not replace missing live evidence with | ||
| a scripted model. | ||
|
|
||
| ## Recorded live execution | ||
|
|
||
| A local paired execution on 2026-10-05 used the CLI commit and version above, Node.js `24.15.0`, provider | ||
| `bigmodel-api`, model `GLM-5.3`, and the Skill pinned to `5cf66be6292eba4ad6e05a4e24e409ebfe1da22a`. | ||
| All 14 user turns completed with native model/tool events and controlled MCP wire evidence. Both arms passed **5/5** | ||
| cases. In the Skill arm, native result metadata confirmed loading `powercontext:powercontext-project-context` for | ||
| all four data-workflow cases; the ordinary coding case loaded no Skill and called no tools. The no-Skill arm loaded | ||
| no PowerContext Skill. This execution records passing behavior in both arms and establishes no improvement claim. | ||
| The full private input/evidence archive supports replay; it is not uploaded by the workflow. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,15 @@ | ||
| # Copyright (c) 2026 OceanBase. | ||
| # | ||
| # Licensed under the Apache License, Version 2.0 (the "License"); | ||
| # you may not use this file except in compliance with the License. | ||
| # You may obtain a copy of the License at | ||
| # | ||
| # http://www.apache.org/licenses/LICENSE-2.0 | ||
| # | ||
| # Unless required by applicable law or agreed to in writing, software | ||
| # distributed under the License is distributed on an "AS IS" BASIS, | ||
| # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. | ||
| # See the License for the specific language governing permissions and | ||
| # limitations under the License. | ||
|
|
||
| """Repository-only evaluation of ZCode Skill routing with native MCP calls.""" |
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
[P2] 新增门禁未覆盖
evaluation.zcode_guidance.run,跨树导入失效不会被任何检查发现本步骤名为「Validate the ZCode Skill pin and behavior evaluation harness」,但只运行
pin --check与pytest evaluation/zcode_guidance/tests。evaluation/zcode_guidance/tests/test_guidance.py:25-27只导入fixture/pin/report,从不导入run;而evaluation/zcode_guidance/run.py:32-33依赖两个跨树符号:tests.e2e.zcode_acceptance.host.NativeHost、tests.e2e.zcode_acceptance.runner.{error_code, serve}。pyproject.toml:193已把evaluation排除出ty check作用域(实测:无参数ty check通过,显式ty check evaluation报 357 条诊断),ruff 也不做跨模块符号解析。因此tests/e2e/zcode_acceptance/内的重命名会让run.py静默失效:既不会被新增的本 workflow 发现,也不会被make quality(Makefile:32的uv run ty check)发现。而run.py正是evaluation/zcode_guidance/README.md中产出 live 报告的入口,report.py的replay()又要求model_mode=live的产物。建议:在本步骤追加一行导入冒烟,例如
uv run --locked --no-sync python -c "import evaluation.zcode_guidance.run",或在
test_guidance.py中加一条importlib.import_module("evaluation.zcode_guidance.run")的测试。