-
Notifications
You must be signed in to change notification settings - Fork 161
579 lines (560 loc) · 34.4 KB
/
Copy pathtest_email_agent_eval.yml
File metadata and controls
579 lines (560 loc) · 34.4 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
# Copyright(C) 2025-2026 Advanced Micro Devices, Inc. All rights reserved.
# SPDX-License-Identifier: MIT
# Offline eval for the email-triage agent (#1112) — weekly, on email PRs,
# on demand, and as a release gate via workflow_call.
#
# Drives the email-triage throughput/quality benchmark over the committed
# SYNTHETIC corpus (tests/fixtures/email/) via FakeGmailBackend — never a live
# mailbox. Computes the quality (FP/FN, #1278) and performance (TTFT / tok-per-s
# / 50-email latency / peak memory, #1277) gates and the categorization-accuracy
# / phishing-precision numbers, logs them, and uploads the full gate report.
#
# SINGLE SOURCE. The gate bars + the single on/off switch live in ONE committed
# source each — the threshold manifests under tests/fixtures/email/, read here
# only through the harness loaders (no thresholds are hardcoded in this YAML):
# - tests/fixtures/email/quality_gate_thresholds.json (FP<5% / FN<2%, #1278)
# - tests/fixtures/email/perf_gate_thresholds.json (TTFT / tps / latency / mem, #1277)
# - tests/fixtures/email/drafting_gate_thresholds.json (draft-approval >=70%, #1269)
# - tests/fixtures/email/briefing_gate_thresholds.json (briefing quality, #1951)
# - tests/fixtures/email/action_items_gate_thresholds.json (P/R/F1, #1949)
#
# ALL FIVE CURRENTLY SHIP `enforce: false` — every gate is in REPORT MODE. #2038
# ("unblock v0.4.0 — calibrate perf bars, report-mode unvalidated judge gates")
# flipped perf, drafting and briefing to false and this header was not updated
# with it; the paragraphs below used to claim they were ENFORCING and were wrong
# from #2038 until this line was written. Read the manifests, not this comment,
# if the two ever disagree again — `git grep '"enforce"' tests/fixtures/email/`.
#
# CI keys off each gate's `should_fail` (= enforce AND not passed). With enforce
# false everywhere, `should_fail` is never true, so a breached BAR cannot fail any
# run on any trigger today. What still fails the build, on every trigger, is the
# class of things that mean the eval could not prove anything: a missing
# ANTHROPIC_API_KEY, an errored or unjudged case, a judge transport error, or a
# harness crash (CLAUDE.md fail-loudly — never an un-judged pass).
#
# Enforcement is toggled IN THE MANIFEST (data, not this workflow): flip a gate's
# `"enforce"` to true and it blocks the weekly, the release AND the PR at once.
# No edit to this file is needed. The bars become meaningful gate-by-gate as
# #1266 (categorization) / #1271 (phishing) land and the judged evals accumulate
# a stable baseline.
#
# NOTE on accuracy gates: categorization >=85% (#1266) and phishing precision
# >=90% (#1271) do not yet have committed threshold manifests (they are owned by
# those issues). This workflow REPORTS the categorization accuracy and phishing
# precision the benchmark already emits, but does not gate on them — there is no
# committed bar to read, and inventing one in YAML would violate the single-source
# rule above. They become gating once their manifests land and this gate-reader is
# pointed at them. Draft-approval >=70% (#1269) DOES have a committed manifest
# (tests/fixtures/email/drafting_gate_thresholds.json) — it is scored by the
# voice-drafting eval step below (#1607 / #1948) and reported, not enforced
# (enforce:false, per #2038: no stable judged baseline yet). A missing judge
# credential still fails the build (eval_drafting_report.py exits 1 — no silent
# skip).
#
# Self-hosted: this needs a running Lemonade Server on AMD hardware, so it runs
# on the [self-hosted, Windows, stx] pool. The install-lemonade action pins the
# runner to the expected LEMONADE_VERSION (installs if missing, reconciles if
# drifted) and ensure-lemonade-running.ps1 starts + warms the server before the
# eval. Serial execution across the eval workflows is enforced by the shared
# `lemonade-eval` concurrency group (now on the JOB — see the note there), NOT
# by a runner label.
#
# ── PULL-REQUEST COVERAGE (#2695 fallout) ─────────────────────────────────────
# Until this trigger existed, a PR could change the email agent's system prompt,
# tools, or triage behavior and merge with ZERO LLM-behavior evidence: pytest and
# the non-LLM integration suites went green and nothing ran an eval. #2695
# (bundled skills, merged as bbf69fd6) is the proof — it injected three Agent
# Skills into the system prompt by default and re-sized the bulk-triage envelope
# budget, having already blown the 16,384-token window once at 16,602 tokens
# during development. It landed at 16,106 — 278 tokens of headroom, established
# by a single manual run and pinned by nothing in CI.
#
# The `pull_request` trigger below closes that. Three things keep it affordable
# on a scarce, strictly-serial runner pool:
#
# 1. PATHS. It fires only on PRs touching the email agent's own code, its
# corpus/threshold manifests, the eval machinery that scores it, the skills
# runtime that injects text into its system prompt, or this workflow and the
# composite actions it runs. Deliberately NOT `src/gaia/agents/base/**` or
# `src/gaia/llm/**`: those are cross-agent surfaces already covered by
# test_eval_agent_gemma_consolidation.yml, and adding them here would roughly
# double the load on the single Lemonade slot. For a core-only change that
# you suspect moves email behavior, `workflow_dispatch` this workflow against
# the branch.
# 2. SUPERSEDING. PR runs share a per-PR concurrency group with
# cancel-in-progress, so rapid pushes never stack up 90-minute jobs.
# 3. LIMIT. PR runs triage 20 emails, not 50 — same slice the release gate uses.
#
# COST TO THE RELEASE PATH, STATED PLAINLY. The `lemonade-eval` group runs one at
# a time and queues the rest FIFO (`queue: max`), so adding ~1 PR run/day costs
# release-gate and weekly calls WAIT time rather than eviction. FIFO has no
# priority lane: a release call sits behind every PR eval already queued, plus any
# in-flight scorecard refresh (up to 14h) or gemma consolidation (up to 7.5h).
# Even so the cost is bounded and visible, not hidden: per-PR cancel-in-progress
# collapses repeated pushes to one run, and weekly_eval.yml treats a non-success
# (not merely a 'failure') callee as a reportable outcome. The residual bad
# outcome is a run that waits out GitHub's 24h job-queue limit and is CANCELLED,
# leaving it with no eval and no auto-requeue — rarer than the evict-the-pending-
# run behavior this replaced, but not gone. Re-run the job from the Checks tab if
# the coverage matters for the change under review. Surfacing that outcome in
# publish.yml / release_agent_email.yml the way weekly_eval.yml already does is
# tracked separately.
#
# A PR RUN IS ADVISORY, AND THAT IS NOT A TRIGGER-SCOPED CHOICE. Per the enforce
# state above, no gate blocks anything on any trigger right now, so a PR run is
# report mode for free. There is deliberately NO "report only on PRs" override
# here: it would break the single-source rule, and — with every manifest at
# enforce:false — the only exits it could soften are the fail-loudly ones. To make
# a gate block a PR, flip its manifest's `"enforce"` to true; it then blocks the
# weekly, the release AND the PR together, which is the right coupling.
#
# WHAT MAKES THE ADVISORY RESULT VISIBLE. A breached bar under enforce:false is a
# GREEN step with no annotation — evidence that exists only inside an artifact
# nobody opens is not coverage. The "Gate verdicts" step near the end of this job
# therefore parses the gate reports the eval writes to eval-out/ and republishes
# every gate's pass/fail into the job summary, raising a `::warning::` per breach.
# That, not the exit code, is the PR-facing signal today.
#
# FORKS. The job does not run for fork PRs — pointing a self-hosted runner at
# unreviewed fork code is not acceptable, and a fork PR carries no
# ANTHROPIC_API_KEY, so it could only ever produce an un-judged result.
# test_eval_agent_gemma_consolidation.yml now triggers on `pull_request` too (a
# `push` filter re-fired it on every main-merge, gating PRs that touched no
# Python), so it carries the same fork and draft clauses in its job `if:` —
# without this workflow's `skip-email-eval` clause. A fork PR touching the email
# agent needs a maintainer to push the branch into this repo and re-run.
name: Email Agent Eval (offline, report mode)
on:
# Weekly — Mondays 07:17 UTC. Was nightly; moved to weekly because the run now
# includes the judge-scored voice-drafting eval, which spends Claude tokens.
# (The release itself re-runs the full suite as a hard gate — see
# release_agent_email.yml eval-gate — so daily cadence here is unnecessary.)
schedule:
- cron: '17 7 * * 1'
# Pre-merge coverage for email-agent behavior changes — see the PULL-REQUEST
# COVERAGE block in the header for why these paths and not more. Advisory
# (report mode) by default; `ready_for_review` is in `types` so undrafting a PR
# triggers the run that the draft state skipped.
pull_request:
branches: [ main ]
types: [opened, synchronize, reopened, ready_for_review]
paths:
# The agent itself — prompts, tools, triage/condense, bundled skills, and
# the packaging/eval_*_report.py gate readers this workflow invokes.
- 'hub/agents/email/python/**'
# The synthetic corpus, ground truth, and every committed threshold
# manifest the gates are read from.
- 'tests/fixtures/email/**'
# The skills runtime injects skill bodies into the agent's system prompt
# and charges them against the triage context budget (#2695).
- 'src/gaia/skills/**'
# The gate machinery this workflow executes: gaia.eval.performance,
# draft_quality, action_item_quality, briefing_quality and the threshold
# loaders. A change here moves the verdict without touching the agent.
- 'src/gaia/eval/**'
- '.github/workflows/test_email_agent_eval.yml'
- '.github/actions/setup-venv/**'
- '.github/actions/install-lemonade/**'
- 'installer/scripts/ensure-lemonade-running.ps1'
# Manual trigger, with a knob for repeat experiments (variance).
workflow_dispatch:
inputs:
model:
description: 'Lemonade model id to benchmark'
required: false
default: 'Gemma-4-E4B-it-GGUF'
limit:
description: 'Max messages to triage from the synthetic corpus'
required: false
default: '50'
experiments:
description: 'Repeat count for variance analysis'
required: false
default: '1'
# Reusable: called as a release gate from the email-agent release
# (release_agent_email.yml) and the GAIA core release (publish.yml) so a
# release runs the same triage/drafting/action-item eval suite it runs
# nightly. The `lemonade-eval` concurrency group below keeps release-triggered
# runs strictly serial with the nightly/refresh runs on the single Lemonade
# slot. Report mode is preserved end-to-end (the manifests own the enforce
# switch), so a call only hard-blocks a release when enforce:true is set.
workflow_call:
inputs:
model:
description: 'Lemonade model id to benchmark'
required: false
type: string
default: 'Gemma-4-E4B-it-GGUF'
limit:
description: 'Max messages to triage from the synthetic corpus'
required: false
type: string
default: '50'
experiments:
description: 'Repeat count for variance analysis'
required: false
type: string
default: '1'
secrets:
ANTHROPIC_API_KEY:
description: 'Judge credential for the drafting / action-item / briefing judged evals — REQUIRED; absent → the preflight step fails the run before any eval spend, never a skip. (Declared optional only so the call site does not have to name it when using `secrets: inherit`.)'
required: false
concurrency:
# PER-INVOCATION SUPERSEDING — NOT the serial slot. The `lemonade-eval` group
# that serializes this against the other self-hosted evals moved to the JOB
# below; it is deliberately NOT repeated here, because a workflow run holding a
# group at the workflow level while its only job waits on the same group at the
# job level is a self-deadlock.
#
# This group exists so a rapid re-push to a PR cancels the 90-minute run it
# superseded instead of queueing a second one behind it.
#
# The non-PR key is `github.run_id`, deliberately, NOT `github.ref`. A called
# reusable workflow's `concurrency` is evaluated in the CALLER's context — this
# repo has the scar to prove it (weekly_eval.yml: a caller holding the group the
# callee needed produced the phantom-job bug, every Weekly Eval run 2026-07-09..19
# dying in <60s and filing issue #2026 on a no-op). So this expression is live on
# the release path too, and keying it to the run id makes each release/weekly run
# the sole member of its own group: it can never queue behind, cancel, or be
# cancelled by anything. A PR key is the PR number so successive pushes to the
# SAME PR share one group and supersede each other.
group: email-agent-eval-${{ github.event.pull_request.number || github.run_id }}
# PR runs only. A weekly/dispatch/release run is never cancelled by a newer one
# (and, being alone in its group, has nothing to be cancelled by).
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
permissions:
contents: read
env:
# `inputs.*` (not `github.event.inputs.*`) so the values propagate for BOTH
# workflow_dispatch AND workflow_call — on a reusable call github.event.inputs
# is empty. On schedule `inputs` is null and each falls back to its default.
EMAIL_EVAL_MODEL: ${{ inputs.model || 'Gemma-4-E4B-it-GGUF' }}
# PR runs triage 20 emails, not 50 — the same slice release_agent_email.yml
# passes. TTFT/throughput are per-token and count-independent, and the three
# judge evals run their own fixed-size seed corpora, so the only thing 20 costs
# is categorization-accuracy sample size. It buys back ~19 min of a scarce,
# strictly-serial runner slot on every email PR.
EMAIL_EVAL_LIMIT: ${{ inputs.limit || (github.event_name == 'pull_request' && '20') || '50' }}
EMAIL_EVAL_EXPERIMENTS: ${{ inputs.experiments || '1' }}
GAIA_MEMORY_DISABLED: "1"
PYTHONIOENCODING: "utf-8"
# Loopback Lemonade — bypass any runner proxy so Python requests reaches it.
# Bare host (no /api/v1): the LemonadeClient normalizes the suffix itself.
LEMONADE_BASE_URL: "http://127.0.0.1:13305"
NO_PROXY: "localhost,127.0.0.1"
# `gaia eval benchmark` is the installed console script; it imports the
# synthetic-corpus backend via `from tests.fixtures.email.fake_gmail import
# FakeGmailBackend`. The repo root must be on sys.path or it dies with
# `ModuleNotFoundError: No module named 'tests'`. Set job-wide so every eval
# step is covered — mirrors email_scorecard_refresh.yml.
PYTHONPATH: ${{ github.workspace }}
# The benchmark triages the whole corpus in ONE `triage_inbox` tool call
# (~38s/email measured on the stx pool for Gemma-4-E4B), so the tool timeout
# must cover the largest limit run here: the weekly default is 50 emails
# (~1900s) and the release gate passes limit=20 (~760s). 1800s abandoned the
# 50-email run mid-flight -> 0 triage results -> Category acc 0.0; 2700s
# (45 min) covers it with margin. (Faster triage is tracked separately.)
GAIA_AGENT_TOOL_TIMEOUT: "2700"
# The agent's calendar-connector resolution blocks on the OS keyring in a
# headless runner — disable it so agent construction/triage doesn't hang.
PYTHON_KEYRING_BACKEND: keyring.backends.null.Keyring
jobs:
email-eval:
name: Email Triage Eval (synthetic corpus, report mode)
# Require the WINDOWS stx runners: the bare 'stx' label also matches a
# Linux box (xsj-aimlab-stxp-04), and this eval is Windows-only (PowerShell
# steps + LemonadeServer.exe) — landing on Linux dies instantly with
# 'powershell: command not found'. fromJSON(...) yields a label ARRAY so the
# runner must have Windows AND the stx/stx-test label.
# The event_name guard keeps the pull_request dereference off schedule /
# workflow_dispatch / workflow_call runs, where that context is absent.
runs-on: ${{ fromJSON(github.event_name == 'pull_request' && contains(github.event.pull_request.labels.*.name, 'stx-test') && '["self-hosted","Windows","stx-test"]' || '["self-hosted","Windows","stx"]') }}
# THE SERIAL SLOT. Same group name as email_scorecard_refresh.yml and
# test_eval_agent_gemma_consolidation.yml — `lemonade-eval`, exactly — so at
# most one eval touches the single Lemonade backend at a time (CLAUDE.md: at
# most one `gaia eval agent` process, period). Two concurrent runs
# race-evict each other's model and produce chaotic ctx/model_load failures.
#
# It lives on the JOB rather than the workflow now, purely so the workflow
# level is free for the per-PR superseding group above. Concurrency groups are
# repository-wide and shared between job-level and workflow-level declarations
# ("only a single job OR WORKFLOW using the same concurrency group will run at
# a time"), so serialization against those sibling workflows is unchanged.
#
# Do NOT also declare `lemonade-eval` at the workflow level of this file: the
# run would hold the group while its only job waited for it — a self-deadlock,
# and the same shape as the phantom-job bug weekly_eval.yml documents.
concurrency:
group: lemonade-eval
cancel-in-progress: false
# `queue: single` (the default) keeps only ONE pending run and cancels the
# previously-queued one — a phantom red X on the PR. `max` queues them FIFO.
queue: max
# PR guard (no-op for schedule / workflow_dispatch / workflow_call):
# - fork PRs never reach the self-hosted pool, and could only produce an
# un-judged result anyway (no ANTHROPIC_API_KEY) — see the header;
# - drafts skip the 90-minute run until marked ready (or labeled
# `ready_for_ci`), matching test_email_agent.yml;
# - `skip-email-eval` is the documented opt-OUT for an email PR that
# provably cannot move behavior (docs, comments, packaging metadata).
# It is opt-OUT on purpose: an opt-IN label would leave the default at
# "no coverage", which is the gap this trigger exists to close.
# Kept flat-indented on purpose: a YAML folded scalar preserves newlines on
# more-indented lines, and an `if:` expression must fold to a single line.
if: >-
github.event_name != 'pull_request' ||
(github.event.pull_request.head.repo.full_name == github.repository &&
(github.event.pull_request.draft == false ||
contains(github.event.pull_request.labels.*.name, 'ready_for_ci')) &&
!contains(github.event.pull_request.labels.*.name, 'skip-email-eval'))
timeout-minutes: 90
steps:
- name: Checkout
uses: actions/checkout@v7
# Fail-fast preflight. The three judged evals (drafting / action-item /
# briefing) each hard-require the judge credential and exit 1 without it —
# but only after the ~40-minute benchmark has already run. Checking it here
# turns "burn 40 minutes of the single serial Lemonade slot, then die" into
# "die in 10 seconds", which matters much more now that PRs queue for that
# slot too. Fails on EVERY trigger — a run that cannot judge cannot produce
# evidence (CLAUDE.md: fail loudly, no silent skip).
# Deliberately NOT a Lemonade check: the "Start Lemonade Server" step below
# is already that preflight and throws when the server is not ready.
- name: Preflight — judge credential
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
# Single-quoted: a backtick in a double-quoted PowerShell string is the
# escape character and would eat the markdown ticks.
if ([string]::IsNullOrWhiteSpace($env:ANTHROPIC_API_KEY)) {
Write-Host '::error::ANTHROPIC_API_KEY is not set. The drafting, action-item and briefing evals score with the Claude judge (src/gaia/eval/claude.py) and exit 1 without it, so this run could only ever produce an un-judged result. Add the ANTHROPIC_API_KEY repository secret, or pass `secrets: inherit` from the calling workflow.'
exit 1
}
Write-Host "Judge credential present."
- name: Setup Python Environment
uses: ./.github/actions/setup-venv
with:
python-version: '3.12'
# [dev] + [eval] for the benchmark + quality/perf gate machinery and
# the Anthropic/scikit-learn judge deps; [api] for the connectors
# layer the email agent boots. The email agent ships as the standalone
# gaia-agent-email wheel (#1102) and is NOT pulled by those extras —
# `gaia eval benchmark` imports gaia_agent_email.agent, so install the
# LOCAL editable hub package (this branch's code, not the PyPI wheel)
# or the benchmark dies with ModuleNotFoundError on a clean venv.
# Core gaia is installed EDITABLE (-e .) too: the eval reads committed
# threshold manifests via gaia.eval's default_*_thresholds_path(), which
# resolve `Path(__file__).parents[3]/tests/fixtures/email/...`. A
# non-editable core lands in site-packages, so parents[3] points at
# .venv/Lib -> FileNotFoundError on the gate manifests. Editable keeps
# __file__ in the repo so the manifests resolve.
install-package: '-e .[dev,eval,api] -e hub/agents/email/python'
- name: Install Lemonade Server
uses: ./.github/actions/install-lemonade
- name: Start Lemonade Server (version-matched, warmed)
timeout-minutes: 15
env:
HUGGINGFACE_ACCESS_TOKEN: ${{ secrets.HUGGINGFACE_ACCESS_TOKEN }}
HF_TOKEN: ${{ secrets.HUGGINGFACE_ACCESS_TOKEN }}
run: |
# Launches the version-matched server as a persistent Scheduled Task
# (survives across CI jobs), reuses it if healthy, and warms the model.
powershell -ExecutionPolicy Bypass -File installer\scripts\ensure-lemonade-running.ps1 -WarmModel "$env:EMAIL_EVAL_MODEL"
if ($LASTEXITCODE -ne 0) { throw "Lemonade server not ready" }
- name: Run email-triage benchmark over the synthetic corpus
run: |
Write-Host "================================================================"
Write-Host " EMAIL TRIAGE EVAL - synthetic corpus (report mode)"
Write-Host "================================================================"
Write-Host "Model: $env:EMAIL_EVAL_MODEL"
Write-Host "Limit: $env:EMAIL_EVAL_LIMIT emails (synthetic inbox)"
Write-Host "Experiments: $env:EMAIL_EVAL_EXPERIMENTS"
Write-Host "Corpus: tests/fixtures/email/synthetic_inbox.mbox (FakeGmailBackend)"
Write-Host "Ground truth: tests/fixtures/email/ground_truth.json"
Write-Host ""
New-Item -ItemType Directory -Force eval-out | Out-Null
# mbox + ground_truth are generated artifacts (not committed) — build
# them from the committed seed before the eval reads them.
python tests/fixtures/email/generate_mbox.py
if ($LASTEXITCODE -ne 0) { throw "generate_mbox failed" }
# Drives EmailTriageAgent over FakeGmailBackend(synthetic_inbox.mbox)
# — synthetic/sandbox only, never a live mailbox. Writes the scorecard
# + variance for the artifact; the gate-reader step below recomputes
# the gates against the committed manifests.
gaia eval benchmark `
--model "$env:EMAIL_EVAL_MODEL" `
--limit "$env:EMAIL_EVAL_LIMIT" `
--experiments "$env:EMAIL_EVAL_EXPERIMENTS" `
--mbox-path tests/fixtures/email/synthetic_inbox.mbox `
--ground-truth tests/fixtures/email/ground_truth.json `
--output-dir eval-out
if ($LASTEXITCODE -ne 0) { throw "gaia eval benchmark failed" }
- name: Evaluate committed gates (report mode — reads the threshold manifests)
run: |
# Gate-reader: reads the gate bars ONLY through the harness loaders
# (single committed thresholds source), runs the gate machinery, logs
# the result, writes eval-out/gate_report.json, and exits non-zero ONLY
# if a gate's should_fail is true (enforce:true AND breached).
python hub/agents/email/python/packaging/eval_gate_report.py
if ($LASTEXITCODE -ne 0) { throw "gate report failed" }
# =================================================================
# BEGIN voice-drafting quality eval (#1607 feature / #1269 metric /
# #1948 tracking) — ADDITIVE, self-contained block. Report mode.
#
# Judge-scored draft quality over the committed drafting seed corpus with
# the #1607 voice profile active (Lemonade, FakeGmailBackend — drafting
# only, nothing is ever sent); a Claude judge scores each draft against
# the case rubric. The aggregate draft_approval_rate is compared to the
# committed manifest tests/fixtures/email/drafting_gate_thresholds.json
# — same single-source rule as the gates above. The manifest currently ships
# enforce:false (#2038), so a breach is REPORTED (see the "Gate verdicts"
# step) rather than failing the build; flip `enforce` there to make it
# block. A missing judge credential fails the run on every trigger (the
# preflight step, and eval_drafting_report.py itself exits 1 — no silent
# skip, CLAUDE.md fail-loudly).
#
# Runs AFTER the triage benchmark in the same job, so Lemonade access stays
# strictly serial (CLAUDE.md: evals serial).
# =================================================================
- name: Voice-drafting quality eval (judge-scored)
env:
# Judge credential for gaia.eval.claude.ClaudeClient — REQUIRED. When the
# secret is absent eval_drafting_report.py exits 1 (fail-loudly, never a
# skip), so this step throws rather than shipping an un-judged gate.
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
python hub/agents/email/python/packaging/eval_drafting_report.py
if ($LASTEXITCODE -ne 0) { throw "drafting eval failed" }
# =================================================================
# END voice-drafting quality eval
# =================================================================
# =================================================================
# BEGIN action-item extraction eval (#1605 feature / #1949 metric /
# #1964 wiring) — ADDITIVE, self-contained block. Report mode.
#
# Precision / recall / F1 of the agent's extracted action items vs a
# hand-labeled corpus (tests/fixtures/email/action_items_ground_truth.json,
# hard negatives included). Generation drives the REAL triage path over a
# FakeGmailBackend (Lemonade — nothing is ever sent). The Claude
# equivalence judge resolves borderline description pairs and is REQUIRED:
# ANTHROPIC_API_KEY MUST be present, and if the judge cannot run the step
# FAILS LOUDLY. There is NO fallback to fuzzy-only matching — a missing or
# broken judge is an error, never a silent degradation to a weaker scorer
# (CLAUDE.md: No Silent Fallbacks — Fail Loudly).
# The aggregate is scored against the committed manifest
# tests/fixtures/email/action_items_gate_thresholds.json (enforce:false)
# via gaia.eval.action_item_quality — same single-source rule as the gates
# above: no thresholds inlined in this YAML; flip `enforce` in the manifest
# (data, not code) to make this gate block once a baseline confirms the bars.
#
# Runs AFTER the triage + drafting evals in the same job, so Lemonade
# access stays strictly serial (CLAUDE.md: evals serial). Kept
# self-contained so concurrent edits to this workflow merge cleanly.
# =================================================================
- name: Action-item extraction eval (report mode)
env:
# Judge credential for the borderline equivalence judge — REQUIRED.
# Absent -> the script fails loudly (NO fuzzy-only fallback).
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
python hub/agents/email/python/packaging/eval_action_item_report.py
if ($LASTEXITCODE -ne 0) { throw "action-item eval failed" }
# =================================================================
# END action-item extraction eval
# =================================================================
# =================================================================
# BEGIN daily-briefing summary-quality eval (#1608 feature / #1951
# tracking) — ADDITIVE, self-contained block. Report mode (was ENFORCING
# until #2038 flipped the manifest).
#
# Judge-scored briefing quality: the REAL scheduled-briefing path
# (gaia_agent_email.briefing.run_briefing_job -> pre_scan_inbox_impl)
# produces the email_pre_scan envelope over the committed briefing seed
# corpus (FakeGmailBackend — read-only, nothing sent/archived), and a
# Claude judge scores each briefing against the case inbox + rubric on
# faithfulness / must-include recall / hallucination-free / grouping. The
# aggregate is compared to the committed manifest
# tests/fixtures/email/briefing_gate_thresholds.json (enforce:false)
# via gaia.eval.briefing_quality — same single-source rule as the gates
# above: no thresholds inlined in this YAML; tune the bars in the manifest
# (data, not code).
#
# NO SILENT SKIP: any errored/unjudged case, a missing judge credential, or
# a judge transport error fails the build on every trigger — if the eval
# cannot prove anything, the pipeline goes red. The QUALITY BAR itself is
# currently reported, not enforced: this manifest was flipped to
# enforce:false by #2038 along with perf and drafting (this comment claimed
# otherwise until the PR trigger landed). Flip `enforce` back to true in the
# manifest to restore blocking; breaches are surfaced meanwhile by the
# "Gate verdicts" step.
#
# Runs AFTER the drafting eval in the same job, so Lemonade access stays
# strictly serial (CLAUDE.md: evals serial). Keep this block
# self-contained so concurrent edits to this workflow merge cleanly.
# =================================================================
- name: Daily-briefing summary-quality eval (judge-scored)
env:
# Judge credential for gaia.eval.claude.ClaudeClient — REQUIRED. This
# gate has no report-mode fallback: an absent key fails the build.
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
python hub/agents/email/python/packaging/eval_briefing_report.py
if ($LASTEXITCODE -ne 0) { throw "briefing eval failed" }
# =================================================================
# END daily-briefing summary-quality eval
# =================================================================
- name: Upload eval scorecard + gate report
if: always()
uses: actions/upload-artifact@v7
with:
name: email-eval-report
path: eval-out/
if-no-files-found: warn
# =================================================================
# GATE VERDICTS — the reporting half of report mode.
#
# Every gate manifest ships enforce:false, so a breached bar exits 0: the eval
# steps are GREEN and the number lives only inside the uploaded artifact. Without
# this step the PR trigger would add a check that cannot say "this got
# worse" — exactly the silent degradation CLAUDE.md forbids. eval_summary.py
# reads the gate reports the eval already wrote and republishes every
# verdict into the job summary, one ::warning:: per breach.
#
# Runs on EVERY trigger, not just PRs: the weekly and the release gate have
# the same blind spot and get the same evidence.
#
# `if: always()` so the verdict is written even when an earlier step
# (preflight / Lemonade / benchmark) failed the job — a reviewer needs to see
# WHICH stage died, and "no gate report was produced" stated out loud beats a
# bare red X. The script carries no gate semantics — it never turns a passing
# run red over a breach, nor rescues a failing one. The `throw` below is not a
# gate either: a nonzero exit means the REPORTER itself broke (missing venv,
# missing file), which is a real failure and must not pass quietly.
# =================================================================
- name: Gate verdicts (job summary)
if: always()
run: |
python hub/agents/email/python/packaging/eval_summary.py eval-out
if ($LASTEXITCODE -ne 0) { throw "eval summary failed" }
- name: Eval summary
if: always()
run: |
Write-Host ""
Write-Host "================================================================"
Write-Host " EMAIL TRIAGE EVAL SUMMARY"
Write-Host "================================================================"
Write-Host "Ran the email-triage benchmark over the committed SYNTHETIC corpus"
Write-Host "(FakeGmailBackend - never a live mailbox) and scored the quality"
Write-Host "(FP/FN) and performance gates from the committed threshold"
Write-Host "manifests under tests/fixtures/email/."
Write-Host ""
Write-Host "ALL gate manifests currently ship enforce:false (report mode), so a"
Write-Host "breached BAR is logged, summarized and uploaded but does NOT fail the"
Write-Host "build - on any trigger, PRs included. See the per-gate verdicts in the"
Write-Host "job summary. Enforcement is toggled in the manifest (data, not this"
Write-Host "workflow) - see the header of this file."
Write-Host ""
Write-Host "What DOES fail the build, everywhere: a missing judge credential, an"
Write-Host "unreachable Lemonade, an unjudged case, or a harness crash. Report mode"
Write-Host "softens 'we measured and it got worse', never 'we did not measure'."
Write-Host "================================================================"