A coding benchmark built around a single task: a TypeScript ledger seeded with intentional bugs. Each model gets the same prompt (see PROMPT.txt). Score each run against hidden tests that the models have no clue about.
npxclaude— Claude Code CLI acting as the harness for all three models- Your Kimi Code key — run
cp .env.example .envand fill it in
Run the benchmark against your model of choice:
./run-fable.sh
./run-opus.sh
./run-kimi.shEach run creates its own runs/<model>-N/ with the patched task code, the model's REPORT.md, and a full transcript.txt. Runs are unattended — permission prompts are bypassed, so the agent executes shell commands in that folder on its own.
To score a finished run, drop the hidden suites into the its dir and run them:
cp hidden/*.test.ts runs/fable-1/
cd runs/fable-1
npx vitest run hidden.defects.test.ts hidden.feature.test.ts