-
-
Notifications
You must be signed in to change notification settings - Fork 83
Expand file tree
/
Copy pathresearch.html
More file actions
307 lines (299 loc) · 18.8 KB
/
Copy pathresearch.html
File metadata and controls
307 lines (299 loc) · 18.8 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
<meta name="description" content="The Citadel research program: open operation control, model-external outcome verification, economic evidence, current proof, and funded milestones." />
<meta name="theme-color" content="#06111f" />
<link rel="canonical" href="https://sethgammon.github.io/Citadel/research.html" />
<link rel="manifest" href="site.webmanifest" />
<meta property="og:type" content="website" />
<meta property="og:site_name" content="Citadel" />
<meta property="og:title" content="Citadel Research Program | Make agent optimization falsifiable" />
<meta property="og:description" content="Open methods, signed receipts, retained failures, prospective local comparisons, and a precise boundary around what Citadel has and has not proved." />
<meta property="og:url" content="https://sethgammon.github.io/Citadel/research.html" />
<meta property="og:image" content="https://sethgammon.github.io/Citadel/assets/citadel-social-preview.png" />
<meta name="twitter:card" content="summary_large_image" />
<meta name="twitter:image" content="https://sethgammon.github.io/Citadel/assets/citadel-social-preview.png" />
<link rel="icon" href="data:image/svg+xml,<svg xmlns=%22http://www.w3.org/2000/svg%22 viewBox=%220 0 32 32%22><rect width=%2232%22 height=%2232%22 rx=%227%22 fill=%22%23070b12%22/><path d=%22M8 8h16v5H13v6h11v5H8z%22 fill=%22%2345ddff%22/></svg>" />
<title>Citadel Research Program | Proof before optimization claims</title>
<link rel="stylesheet" href="site-system.css?v=20260803-1" />
</head>
<body class="site-page site-proof site-research">
<a class="site-skip-link" href="#main-content">Skip to content</a>
<div class="site-scroll-progress" aria-hidden="true"><span></span></div>
<nav class="site-nav" aria-label="Primary navigation">
<div class="site-nav-inner">
<a class="site-brand" href="index.html"><span class="site-brand-mark">C</span><span>Citadel</span></a>
<div class="site-nav-links">
<a class="site-nav-link" href="index.html#product-story">How it works</a>
<a class="site-nav-link" href="evidence.html">Evidence</a>
<a class="site-nav-link" href="operation-control.html">Operation Control</a>
<a class="site-nav-link" href="optimizer.html">Optimizer</a>
<a class="site-nav-link" href="research.html" aria-current="page">Research</a>
</div>
<div class="site-nav-actions">
<a class="site-nav-github" href="https://github.com/SethGammon/Citadel">GitHub</a>
<a class="site-nav-cta" href="#milestones">Funded work</a>
<button class="site-nav-toggle" type="button" data-site-nav-toggle aria-expanded="false" aria-controls="site-nav-mobile">Menu</button>
</div>
</div>
<div class="site-nav-mobile" id="site-nav-mobile" data-site-nav-mobile hidden>
<a class="site-nav-link" href="index.html">Overview</a>
<a class="site-nav-link" href="evidence.html">Evidence index</a>
<a class="site-nav-link" href="operation-control.html">Operation Control</a>
<a class="site-nav-link" href="optimizer.html">Optimizer</a>
<a class="site-nav-link" href="#evidence">Evidence ledger</a>
<a class="site-nav-link" href="#milestones">Funded milestones</a>
<a class="site-nav-link" href="#claims">Claim boundary</a>
</div>
</nav>
<main class="shell" id="main-content" tabindex="-1">
<header class="hero">
<div>
<div class="eyebrow">Open agent optimization research</div>
<h1>Make agent optimization <span>falsifiable.</span></h1>
<p class="hero-copy">
Citadel is building the evidence layer between an optimization policy
and its claim. It controls what an agent operation may run, records
what actually ran, and asks a deterministic verifier outside the routed model whether the work
counts. <strong>A cheaper failure is not a saving. An unknown is not zero.</strong>
</p>
<div class="site-button-row" style="margin-top:30px;">
<a class="site-button primary" href="#evidence">Inspect the evidence</a>
<a class="site-button" href="https://github.com/SethGammon/Citadel/blob/main/docs/grants/SENTIENT_OPTIMIZER_APPLICATION_DRAFT.md">Read the grant draft</a>
</div>
</div>
<aside class="claim-card" aria-label="Research status">
<div class="claim-label">Research status</div>
<h2>The measurement loop works. The outside baseline did not.</h2>
<p>
A calibrated synthetic hybrid preserved 12/12 verified operations at
38.7% lower comparison cost. The later outside-authored diagnostic
verified only 3/16 for the controller and 2/16 for direct Claude, so
Citadel rejects a general optimization claim.
</p>
<code>outside controller / Claude: 3 / 2 of 16<br />comparison cost: -1.26%; not actual cash<br />general production result: not demonstrated</code>
</aside>
</header>
<div class="proof-strip" aria-label="Citadel research evidence summary">
<div class="proof-stat"><strong>240</strong><span>signed prospective comparison cells</span></div>
<div class="proof-stat"><strong>24</strong><span>outside-authored repositories</span></div>
<div class="proof-stat"><strong>32/32</strong><span>official holdout verdicts</span></div>
<div class="proof-stat"><strong>38.7%</strong><span>bounded synthetic reduction</span></div>
</div>
<section id="thesis">
<div class="section-heading">
<div><div class="proof-kicker">The research thesis</div><h2>Optimize the whole operation, then prove the outcome.</h2></div>
<p>
Prompt routing is only one decision. Real agent economics also depend
on topology, decomposition, retries, tools, timeouts, local versus
hosted execution, verification, and recovery. Citadel turns those
choices into a bounded contract whose result can be reproduced and
rejected.
</p>
</div>
<div class="research-thesis" data-site-reveal>
<article class="research-thesis-card">
<div class="proof-kicker">What is different</div>
<h3>The optimizer does not grade its own homework.</h3>
<p>
A selected executor cannot become the winner because it sounded
confident, reported a low number, or produced a patch. Requested
and observed runtime facts are reconciled, required artifacts are
checked, and a model-external repository verifier decides completion.
</p>
</article>
<article class="research-thesis-card target">
<div class="proof-kicker">Funded target</div>
<h3>A result strong enough to survive a no.</h3>
<p>Across a preregistered multi-stack benchmark, the funded controller targets:</p>
<div class="target-pair">
<div class="target-number"><strong>≥80%</strong><span>absolute verified completion</span></div>
<div class="target-number"><strong>≥95%</strong><span>of a valid frontier baseline</span></div>
<div class="target-number"><strong>≥30%</strong><span>lower measured end-to-end cost</span></div>
</div>
<p style="margin-top:16px;font-size:13px;">Frontier must first verify at least 80% overall and 70% in every frozen task stratum, or the comparison is baseline-invalid.</p>
</article>
</div>
</section>
<section id="evidence">
<div class="section-heading">
<div><div class="proof-kicker">Evidence ladder</div><h2>What exists today, in order of claim strength.</h2></div>
<p>
Each rung answers a different question. Implementation evidence shows
the mechanism exists. Retrospective evidence calibrates it. Prospective
evidence shows the public runtime seam works. Six comparative studies
show the progression from timeout sensitivity and escalation regressions
to a passed, narrowly calibrated hybrid support envelope.
</p>
</div>
<div class="proof-ledger">
<article class="proof-ledger-item">
<div class="ledger-id">01 / PRODUCT</div>
<h3>Durable operating layer</h3>
<p><code>/do</code>, repository state, campaigns, fleets, recovery, evidence, and handoffs work around the coding agent a developer already uses.</p>
<span class="ledger-state">Implemented</span>
</article>
<article class="proof-ledger-item">
<div class="ledger-id">02 / HISTORY</div>
<h3>Signed 120-cell matrix</h3>
<p>Ten frozen scenarios across three repositories and four economic policies preserve 33 verified, 51 failed, and 36 unknown outcomes.</p>
<span class="ledger-state">Verified</span>
</article>
<article class="proof-ledger-item">
<div class="ledger-id">03 / STACK</div>
<h3>Sentient ROMA binding</h3>
<p>A thin adapter controlled a pinned recursive solver module by module. Its 24-cell diagnostic passed the evidence gate and failed the efficiency hypothesis.</p>
<span class="ledger-state">Verified</span>
</article>
<article class="proof-ledger-item">
<div class="ledger-id">04 / RUNTIME</div>
<h3>Prospective public task</h3>
<p>One preregistered Claude Code operation on a fresh public clone matched model and topology, changed the required artifact, and passed a deterministic verifier outside the model.</p>
<span class="ledger-state">Verified</span>
</article>
<article class="proof-ledger-item">
<div class="ledger-id">05 / ECONOMICS V1</div>
<h3>Adaptive local comparison</h3>
<p>Across 12 tasks × 2 policies × 3 timing repetitions, adaptive recorded 27/36 verified cells versus 24/36. The frozen aggregate used 9.9% less GPU energy, but excluding one same-route timeout pair reverses the comparison to 3.5% more.</p>
<span class="ledger-state">Verified negative gate</span>
</article>
<article class="proof-ledger-item">
<div class="ledger-id">06 / ECONOMICS V2</div>
<h3>Capability-profile falsification</h3>
<p>A separately frozen 72-cell follow-up matched 24/36 baseline cell completion, but 12 escalations caused 15.7% more GPU energy and 16.4% more modeled GPU cost.</p>
<span class="ledger-state">Verified regression</span>
</article>
<article class="proof-ledger-item">
<div class="ledger-id">07 / REPOSITORY</div>
<h3>Representative fixture shakedown</h3>
<p>Six artifact-producing fixture tasks, two policies, and two timing repetitions produced 24 signed cells. Both policies verified 6/12; zero false passes and path violations survived replay, while the 7.1% energy reduction missed the frozen 20% gate.</p>
<span class="ledger-state">Integrity passed; economics failed</span>
</article>
<article class="proof-ledger-item">
<div class="ledger-id">08 / HYBRID CALIBRATION</div>
<h3>Valid baseline, narrow miss</h3>
<p>Claude and Citadel each verified 12/12 fresh tasks. Citadel avoided four Claude calls and reduced comparison cost 28.4%, missing the frozen 30% gate.</p>
<span class="ledger-state">Quality passed; economics failed</span>
</article>
<article class="proof-ledger-item">
<div class="ledger-id">09 / HYBRID V2</div>
<h3>Calibrated support envelope</h3>
<p>On twelve new tasks, both policies verified 12/12. Citadel used local 3B eight times, recovered once, reduced Claude calls from twelve to five, and reduced comparison cost 38.7%.</p>
<span class="ledger-state">Every frozen gate passed</span>
</article>
<article class="proof-ledger-item">
<div class="ledger-id">10 / PUBLIC HOLDOUT</div>
<h3>Outside-authored baseline falsification</h3>
<p>Twenty-four distinct repositories produced eight calibration and sixteen untouched evaluation tasks. Direct Claude verified 2/16 and the sealed controller 3/16 at 1.26% lower comparison cost. The baseline was too weak for a general result.</p>
<span class="ledger-state">Evidence complete; baseline invalid</span>
</article>
</div>
</section>
<section id="milestones">
<div class="section-heading">
<div><div class="proof-kicker">Funded work</div><h2>Four milestones, each with a failure condition.</h2></div>
<p>
The grant does not fund a promise to make Citadel look intelligent.
It funds a public research program whose method, artifacts, cost
lenses, and negative results remain inspectable whether the final
performance target passes or fails.
</p>
</div>
<div class="milestone-grid">
<article class="milestone-card" data-step="01">
<small>Controller gate</small>
<h3>Learn operation value</h3>
<p>Choose models, topology, decomposition, retries, tools, and stopping from calibrated outcome evidence instead of prompt difficulty alone.</p>
<div class="milestone-gate">Exit: held-out decisions are reproducible and every rejected or unknown outcome remains in the ledger.</div>
</article>
<article class="milestone-card" data-step="02">
<small>Economic gate</small>
<h3>Measure the complete cost</h3>
<p>Separate reported, derived, market-equivalent, marginal, and end-to-end cost. Include local hardware and energy without converting missing evidence to zero.</p>
<div class="milestone-gate">Exit: every published comparison carries a complete cost basis or is explicitly blocked.</div>
</article>
<article class="milestone-card" data-step="03">
<small>Portability gate</small>
<h3>Generalize the adapter</h3>
<p>Exercise the same operation contract across multiple open and proprietary agent stacks without replacing their native planning or execution logic.</p>
<div class="milestone-gate">Exit: requested versus observed identity, artifacts, and outcome receipts reconcile across each stack.</div>
</article>
<article class="milestone-card" data-step="04">
<small>Scale gate</small>
<h3>Run the prospective study</h3>
<p>Extend the three local studies to multiple stacks, repositories, model families, hardware profiles, and tool routes. Freeze every task, baseline, policy, verifier, stopping rule, and economic target before execution.</p>
<div class="milestone-gate">Exit: the result reproduces offline and in clean hosted verification, even if the performance hypothesis fails.</div>
</article>
</div>
</section>
<section id="claims">
<div class="section-heading">
<div><div class="proof-kicker">Claim boundary</div><h2>Credibility comes from saying exactly where the proof ends.</h2></div>
<p>
Citadel has a positive synthetic result and a later outside-authored
diagnostic that invalidated its baseline. The controller can be held
accountable when a policy or comparison fails. It does not yet have
the complete-cost, strong-baseline, multi-stack evidence needed for a
general claim.
</p>
</div>
<div class="claim-boundary-list">
<article class="claim-boundary-column">
<h3>Demonstrated</h3>
<ul>
<li>Stack-neutral operation contracts can bind real agent runtimes.</li>
<li>Requested and observed model identity can be reconciled without inference.</li>
<li>Model-external deterministic verification can reject false, partial, and adversarial outputs.</li>
<li>Signed evidence can preserve passed, failed, and unknown outcomes.</li>
<li>Clean hosted verification reproduces the committed proof bundle.</li>
<li>Four retained calibration studies reject timeout-sensitive, escalation-heavy, baseline-invalid, and narrowly missed policies.</li>
<li>A separately frozen hybrid v2 preserved 12/12 completions and reduced comparison cost 38.7% inside a preregistered support envelope.</li>
<li>A later pilot sealed routes across 24 outside-authored repositories and published all 32 official evaluation verdicts.</li>
<li>The outside-authored result rejected a general claim because direct Claude verified only 2/16.</li>
</ul>
</article>
<article class="claim-boundary-column">
<h3>Not demonstrated</h3>
<ul>
<li>Best-in-class agent performance or broad benchmark leadership.</li>
<li>A general quality advantage across repositories or model families.</li>
<li>Lower latency than direct or frontier-only execution.</li>
<li>Thirty percent lower actual end-to-end cash across outside-authored production tasks with a valid strong baseline.</li>
<li>Production reliability across many external users and environments.</li>
</ul>
</article>
</div>
</section>
<section class="closing">
<article class="closing-card">
<div class="eyebrow">Why Sentient</div>
<h2>A native fit for token and economic optimization.</h2>
<p>
Sentient's product request asks for a layer that can optimize agents
across models and workflows. Citadel contributes the operation-level
control and proof discipline needed to tell a genuine economic win
from a cheaper failed attempt. The first ROMA adapter makes that fit
concrete rather than hypothetical.
</p>
<div class="links">
<a class="button primary" href="https://github.com/SethGammon/Citadel/tree/main/benchmarks/hybrid-economic-pilot-v2">Inspect passed hybrid evidence</a>
<a class="button" href="operation-control.html">Operation proof</a>
<a class="button" href="optimizer.html">Optimizer proof</a>
<a class="button" href="https://sentient.foundation/product-requests">Sentient request</a>
</div>
</article>
<aside class="closing-card">
<div class="claim-label">Public research principle</div>
<p style="font-size:22px;line-height:1.45;color:var(--site-text);">A failed frozen hypothesis is a result. A missing receipt is not.</p>
<p style="font:11px/1.65 var(--site-mono);color:var(--site-dim);">Observed 2026-08-01 · application not yet submitted</p>
</aside>
</section>
</main>
<footer>
<div class="shell">Citadel research program · operation control, model-external verification, and honest agent economics · bounded performance gate passed; generalization gate open.</div>
</footer>
<script src="site-system.js?v=20260803-1"></script>
</body>
</html>