| Result | Pass@1 | Multi turn with feedback (turn=2) | Trials | Samples |
|---|---|---|---|---|
| GLM-4.7-Flash base | 0.5/26 mean | 4.5/26 mean | 4 | 104 |
| Luna | 6.25/26 mean | 16.75/26 mean | 4 | 104 |
| Result | Pass@1 SD, range, 95% CI (out of 26) | Multi turn with feedback (turn=2) SD, range, 95% CI (out of 26) | Conditional turn-2 recovery |
|---|---|---|---|
| GLM-4.7-Flash base | 0.58; 0-1; 0-1.25 | 1.29; 3-6; 2.25-6.75 | 16/102 (15.7%; CI 7.8-24.8%) |
| Luna | 1.89; 5-9; 3.25-9.5 | 1.5; 15-18; 12.75-20.5 | 42/79 (53.2%; CI 36.8-70%) |
| Result | Pass@1 | Multi turn with feedback (turn=2) | Trials | Samples |
|---|---|---|---|---|
| SFT v5, Aider-format | 6/26 mean | 10.25/26 mean | 4 | 104 |
| Synth v1, epoch 50 | 9.5/26 mean | 12/26 mean | 4 | 104 |
| execution-midband-RL-v1 | 8.25/26 mean | 13/26 mean | 4 | 104 |
| execution-midband-RL-v2 | 10.5/26 mean | 14.5/26 mean | 4 | 104 |
| Phone Number kernel12 GRPO20, iter 14 | 11.25/26 mean | 15.25/26 mean | 4 | 104 |
| Generalized C++ kernel GRPO20, iter 14 | 11.75/26 mean | 16/26 mean | 4 | 104 |
| Result | Pass@1 SD, range, 95% CI (out of 26) | Multi turn with feedback (turn=2) SD, range, 95% CI (out of 26) | Conditional turn-2 recovery |
|---|---|---|---|
| SFT v5, Aider-format | 1.63; 4-8; 3-9.25 | 1.71; 8-12; 7-13.75 | 17/80 (21.2%; CI 11.9-31.6%) |
| Synth v1, epoch 50 | 0.58; 9-10; 5.75-13.5 | 0.82; 11-13; 8.25-15.75 | 10/66 (15.2%; CI 7.4-24.4%) |
| execution-midband-RL-v1 | 2.06; 6-11; 4.75-12 | 0.82; 12-14; 9-17 | 19/71 (26.8%; CI 14.5-41.4%) |
| execution-midband-RL-v2 | 1.29; 9-12; 7-14 | 2.08; 12-17; 10.5-18.5 | 16/62 (25.8%; CI 12.9-41.9%) |
| Phone Number kernel12 GRPO20, iter 14 | 0.5; 11-12; 7.5-15 | 0.5; 15-16; 11-19.25 | 16/59 (27.1%; CI 13.1-44.2%) |
| Generalized C++ kernel GRPO20, iter 14 | 1.50; 10-13; 8-15.5 | 1.41; 14-17; 11.75-20 | 17/57 (29.8%; CI 14.5-48.9%) |
Statistics: machine summary · per-task frequencies · recompute
SFT v5 artifacts: checkpoint · training dataset · evaluation evidence
Synth v1 artifacts: reproducibility bundle · checkpoint archive · training dataset · evaluation archive · W&B run
The evaluation archives require authorized WootzappLab HF access and remain evaluation-only. Their payloads are unchanged; historical result manifests retain their source identities. The SFT v5 training-dataset URL records historical provenance: its URL/hash mapping remains unverified and must not be treated as a verified current download.
execution-midband-RL-v1 artifacts: run archive · final adapter · evaluation evidence · evaluation method
execution-midband-RL-v2 artifacts: run archive · final adapter · evaluation evidence · evaluation method · launch configurations · W&B run
Phone Number kernel12 GRPO20 artifacts: run archive · scored adapter · training dataset · evaluation evidence · evaluation method · launch configurations · W&B run
Generalized C++ kernel GRPO20 artifacts: run archive · scored adapter, iter 14 · training dataset · evaluation evidence · W&B run
Generalized C++ result boundary: this row uses four selected, receipt-verified fixed26-contract-v2 trials. It is an assisted regression result, not a random four-trial or pristine held-out benchmark claim; six training task IDs overlap Fixed26.
cd results/base-fixed26-20260711/reproduction
export OPENAI_API_BASE=http://127.0.0.1:8000/v1
export OPENAI_API_KEY=local-eval
./run.shcd results/luna-fixed26-20260805/reproduction
export OPENROUTER_API_KEY=...
./run.shsky launch -y -c fixed26-mt2-v2 results/execution-midband-rl-v2/method/skypilot-task.yamlsky jobs launch -y \
--env EVAL_RUN_ID=<fresh-fixed26-run-id> \
results/phone-number-kernel12-GRPO20/launch-configs/fixed26-eval.yaml