Skip to content

Repository files navigation

Fixed26 results

Reference evaluations

Result Pass@1 Multi turn with feedback (turn=2) Trials Samples
GLM-4.7-Flash base 0.5/26 mean 4.5/26 mean 4 104
Luna 6.25/26 mean 16.75/26 mean 4 104
Result Pass@1 SD, range, 95% CI (out of 26) Multi turn with feedback (turn=2) SD, range, 95% CI (out of 26) Conditional turn-2 recovery
GLM-4.7-Flash base 0.58; 0-1; 0-1.25 1.29; 3-6; 2.25-6.75 16/102 (15.7%; CI 7.8-24.8%)
Luna 1.89; 5-9; 3.25-9.5 1.5; 15-18; 12.75-20.5 42/79 (53.2%; CI 36.8-70%)

Post-training evaluations

Result Pass@1 Multi turn with feedback (turn=2) Trials Samples
SFT v5, Aider-format 6/26 mean 10.25/26 mean 4 104
Synth v1, epoch 50 9.5/26 mean 12/26 mean 4 104
execution-midband-RL-v1 8.25/26 mean 13/26 mean 4 104
execution-midband-RL-v2 10.5/26 mean 14.5/26 mean 4 104
Phone Number kernel12 GRPO20, iter 14 11.25/26 mean 15.25/26 mean 4 104
Generalized C++ kernel GRPO20, iter 14 11.75/26 mean 16/26 mean 4 104
Result Pass@1 SD, range, 95% CI (out of 26) Multi turn with feedback (turn=2) SD, range, 95% CI (out of 26) Conditional turn-2 recovery
SFT v5, Aider-format 1.63; 4-8; 3-9.25 1.71; 8-12; 7-13.75 17/80 (21.2%; CI 11.9-31.6%)
Synth v1, epoch 50 0.58; 9-10; 5.75-13.5 0.82; 11-13; 8.25-15.75 10/66 (15.2%; CI 7.4-24.4%)
execution-midband-RL-v1 2.06; 6-11; 4.75-12 0.82; 12-14; 9-17 19/71 (26.8%; CI 14.5-41.4%)
execution-midband-RL-v2 1.29; 9-12; 7-14 2.08; 12-17; 10.5-18.5 16/62 (25.8%; CI 12.9-41.9%)
Phone Number kernel12 GRPO20, iter 14 0.5; 11-12; 7.5-15 0.5; 15-16; 11-19.25 16/59 (27.1%; CI 13.1-44.2%)
Generalized C++ kernel GRPO20, iter 14 1.50; 10-13; 8-15.5 1.41; 14-17; 11.75-20 17/57 (29.8%; CI 14.5-48.9%)

Statistics: machine summary · per-task frequencies · recompute

SFT v5 artifacts: checkpoint · training dataset · evaluation evidence

Synth v1 artifacts: reproducibility bundle · checkpoint archive · training dataset · evaluation archive · W&B run

The evaluation archives require authorized WootzappLab HF access and remain evaluation-only. Their payloads are unchanged; historical result manifests retain their source identities. The SFT v5 training-dataset URL records historical provenance: its URL/hash mapping remains unverified and must not be treated as a verified current download.

execution-midband-RL-v1 artifacts: run archive · final adapter · evaluation evidence · evaluation method

execution-midband-RL-v2 artifacts: run archive · final adapter · evaluation evidence · evaluation method · launch configurations · W&B run

Phone Number kernel12 GRPO20 artifacts: run archive · scored adapter · training dataset · evaluation evidence · evaluation method · launch configurations · W&B run

Generalized C++ kernel GRPO20 artifacts: run archive · scored adapter, iter 14 · training dataset · evaluation evidence · W&B run

Generalized C++ result boundary: this row uses four selected, receipt-verified fixed26-contract-v2 trials. It is an assisted regression result, not a random four-trial or pristine held-out benchmark claim; six training task IDs overlap Fixed26.

How to reproduce

cd results/base-fixed26-20260711/reproduction
export OPENAI_API_BASE=http://127.0.0.1:8000/v1
export OPENAI_API_KEY=local-eval
./run.sh
cd results/luna-fixed26-20260805/reproduction
export OPENROUTER_API_KEY=...
./run.sh
sky launch -y -c fixed26-mt2-v2 results/execution-midband-rl-v2/method/skypilot-task.yaml
sky jobs launch -y \
  --env EVAL_RUN_ID=<fresh-fixed26-run-id> \
  results/phone-number-kernel12-GRPO20/launch-configs/fixed26-eval.yaml

About

No description, website, or topics provided.

Resources

Code of conduct

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages