Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Kimi K3 vs Claude Fable 5 vs Claude Opus 4.8

A coding benchmark built around a single task: a TypeScript ledger seeded with intentional bugs. Each model gets the same prompt (see PROMPT.txt). Score each run against hidden tests that the models have no clue about.

Requirements

  • npx
  • claude — Claude Code CLI acting as the harness for all three models
  • Your Kimi Code key — run cp .env.example .env and fill it in

Run

Run the benchmark against your model of choice:

./run-fable.sh
./run-opus.sh
./run-kimi.sh

Each run creates its own runs/<model>-N/ with the patched task code, the model's REPORT.md, and a full transcript.txt. Runs are unattended — permission prompts are bypassed, so the agent executes shell commands in that folder on its own.

To score a finished run, drop the hidden suites into the its dir and run them:

cp hidden/*.test.ts runs/fable-1/
cd runs/fable-1
npx vitest run hidden.defects.test.ts hidden.feature.test.ts

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages