Skip to content

feat(evaluation): add pinned Claude Code Skill guidance regression suite - #1748

Merged
Teingi merged 3 commits into
oceanbase:masterfrom
Inference1:feat/1725-skill-guidance-evaluation
Sep 30, 2026
Merged

Teingi merged 3 commits into
oceanbase:masterfrom
Inference1:feat/1725-skill-guidance-evaluation

Conversation

@Inference1

@Inference1 Inference1 commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Which issue or RFC does this PR close?

Closes #1725.

Rationale for this change

Skill routing regressions can pass vacuously when a negative assertion uses the wrong MCP name or the model never calls any tool. Add paired positive/negative controls tied to a pinned packaged Skill and raw Claude session evidence.

What changes are included in this PR?

  • Add a skill-up v0.12.0 Claude Code suite for ordinary coding, explicit saves, empty search, candidate inspection and failed saves, with all 34 mocked operations and shared neutral parameter signatures.
  • Pin/vendor the unchanged packaged Skill; check drift and save exact input snapshots with each run.
  • Retain both arms and native JSON/JUnit/HTML plus raw JSONL; report exact tool names, case-level deltas and evidence completeness separately from behavior failures. Check every relevant call's Scope, matching tool results and a completed assistant response for every turn.
  • Add a credential-free CI validation job and complementary English/Chinese documentation.
  • Keep all 202 assertions while compacting repetitive YAML and removing general validation duplicated by the runner.

Are there any user-facing changes?

New evaluation commands and reports only. No runtime APIs or Skill prose change. Mocked MCP and the controlled Scope helper do not qualify real persistence, authentication, host approval, automatic Capture/Flush, bounded recall or memory quality. with_skill means installed/available, not proof of full Skill-body consumption.

How was this change tested?

  • skill-up validate/dry-run, pin/fixture validation: passed.
  • 26 unittest regressions, Ruff lint/format and local patch checks: passed.
  • Claude Code 2.1.283 through PackyAPI, requested claude-sonnet-4-6: with_skill 5/5, without_skill 5/5, delta +0 percentage points, gate PASS, complete session evidence True.
  • Complete evaluation evidence: native JSON/JUnit/HTML, raw session JSONL, exact input snapshots, gateway metadata and checksums. packyapi-20260926-run4 is the final run; run3 is retained diagnostic evidence from before the shared parameter reference and formatting fixes. Gateway configuration is not independent model identity verification.
  • Rechecked the retained Run4 sessions with the updated reporter: both arms remain 5/5, evidence complete. On separate copies, an extra wrong-Scope write and a transcript truncated at line 19 both passed the old reporter and fail the updated reporter. Original evidence hashes are unchanged; no new model run was needed because parsed cases, prompts, fixtures and Skill inputs are unchanged.
  • uv run prek run -a: all 10 non-type hooks passed; native Windows ty failed on unchanged POSIX code (fcntl, resource, os.O_NOFOLLOW, os.mkfifo). uv run --locked ty check --python-platform linux passed.
  • Website pnpm check:links passed after checking 225 public pages and 332 repository files; links from generated docs to the evaluation project use the upstream GitHub URL.
  • All 24 GitHub CI checks passed at 2c2d727, including Python 3.11–3.14, Windows portability and SQLite/OceanBase acceptance.

An independent audit of all 10 raw sessions confirmed the shared contract, complete 34-tool catalogs, required calls,
arguments and denied-write response. Only the installed arm advertised the Skill; no full-body consumption was observed,
so the zero delta does not establish a Skill benefit. The full repository test suite was not run locally; GitHub CI status is reported by this PR's checks.

AI usage statement

OpenAI Codex (GPT-6) assisted implementation and review. Claude Code model invocations supplied the explicitly identified evaluation data.

@Inference1

Inference1 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

I am very interested in powercontext and am a power user, so I am willing to contribute whenever I have time. Are there any other review suggestions

@Teingi
Teingi merged commit 3affafa into oceanbase:master Sep 30, 2026
25 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(evaluation): evaluate the integration Skill guidance with skill-up

1 participant