Goodhart-proof AI coding pipeline with architectural isolation
-
Updated
Jun 5, 2026 - Python
Goodhart-proof AI coding pipeline with architectural isolation
Standing Algebra (Σᴿ): A Closure-Theoretic Operator for Constraining Domination and Preserving Autonomy
Self-improvement loop for LLM agents with an integrity gate that catches reward hacking: reward gains that erode reproducibility get reverted, even ones a reward-only gate would accept.
Does a CLAUDE.md actually change how Claude behaves? An ablation harness: run adversarial traps with the rules and without them, grade blind, and test whether the difference is real.
Catch reward traps before training. Static analysis for RL reward functions.
The forge, distilled: an ontology of three weeks of alignment research — every direction tried, colored verified / falsified / open, each color backed by a named artifact. Products: justitia, proxylimen, fallacy-cutter. Full tree at tag forge-full-tree.
A preregistered synthetic simulation of proxy failure under optimization pressure, with deterministic replay, matched controls, and adversarial verification.
Toy 6. An interactive phase-space instrument mapping Ψ = S/D — the ratio of capability to modeling depth that determines whether a system is in the viable, transitional, or failure-mode-dominant regime. Includes the Inner Crossing animation. Companion simulation for The Inner Crossing — Series 2, Part 3.
Toy 5. An interactive proxy decay simulator showing how optimization pressure erodes the modeling capacity required to distinguish proxy from territory — producing self-reinforcing V(t) degradation that becomes progressively harder to correct. Companion simulation for The Depth Constraint — Series 2, Part 2.
Chain-of-density study of 'Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents' (Zhang et al. 2026, arXiv:2607.12790) - five-tier note, locator-verified. Double Ratchet, drawback detectors, anchor discipline, Goodhart repair.
Self-Optimizing System Loop — autonomous software optimization via Claude Code. Point an AI agent at a metric. Go to sleep. Wake up with improvements.
Add a description, image, and links to the goodharts-law topic page so that developers can more easily learn about it.
To associate your repository with the goodharts-law topic, visit your repo's landing page and select "manage topics."