Repository navigation
feat(evaluation): add pinned Claude Code Skill guidance regression suite - #1748
Merged
Teingi merged 3 commits intoSep 30, 2026
Merged
Conversation
Inference1
requested review from
AlexStocks,
PsiACE,
frf12 and
zhanghuidinah
as code owners
September 27, 2026 03:59
Contributor
Author
|
I am very interested in |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue or RFC does this PR close?
Closes #1725.
Rationale for this change
Skill routing regressions can pass vacuously when a negative assertion uses the wrong MCP name or the model never calls any tool. Add paired positive/negative controls tied to a pinned packaged Skill and raw Claude session evidence.
What changes are included in this PR?
Are there any user-facing changes?
New evaluation commands and reports only. No runtime APIs or Skill prose change. Mocked MCP and the controlled Scope helper do not qualify real persistence, authentication, host approval, automatic Capture/Flush, bounded recall or memory quality.
with_skillmeans installed/available, not proof of full Skill-body consumption.How was this change tested?
claude-sonnet-4-6: with_skill 5/5, without_skill 5/5, delta +0 percentage points, gate PASS, complete session evidence True.packyapi-20260926-run4is the final run;run3is retained diagnostic evidence from before the shared parameter reference and formatting fixes. Gateway configuration is not independent model identity verification.uv run prek run -a: all 10 non-type hooks passed; native Windowstyfailed on unchanged POSIX code (fcntl,resource,os.O_NOFOLLOW,os.mkfifo).uv run --locked ty check --python-platform linuxpassed.pnpm check:linkspassed after checking 225 public pages and 332 repository files; links from generated docs to the evaluation project use the upstream GitHub URL.2c2d727, including Python 3.11–3.14, Windows portability and SQLite/OceanBase acceptance.An independent audit of all 10 raw sessions confirmed the shared contract, complete 34-tool catalogs, required calls,
arguments and denied-write response. Only the installed arm advertised the Skill; no full-body consumption was observed,
so the zero delta does not establish a Skill benefit. The full repository test suite was not run locally; GitHub CI status is reported by this PR's checks.
AI usage statement
OpenAI Codex (GPT-6) assisted implementation and review. Claude Code model invocations supplied the explicitly identified evaluation data.