feat(daemon): start and supervise the local model server #260
Triggered via pull request
September 3, 2026 03:14
Status
Failure
Total duration
1h 45m 28s
Artifacts
1
test_email_agent_eval.yml
on: pull_request
Email Triage Eval (synthetic corpus, report mode)
31m 35s
Annotations
1 error and 5 warnings
|
Email Triage Eval (synthetic corpus, report mode)
Process completed with exit code 1.
|
|
Email Triage Eval (synthetic corpus, report mode)
Email eval - Daily-briefing quality produced no verdict at all - its eval step did not get far enough to write a report. Read it as missing evidence, not as a pass.
|
|
Email Triage Eval (synthetic corpus, report mode)
Email eval - Action-item extraction produced no verdict at all - its eval step did not get far enough to write a report. Read it as missing evidence, not as a pass.
|
|
Email Triage Eval (synthetic corpus, report mode)
Email eval - Voice-drafting approval produced no verdict at all - its eval step did not get far enough to write a report. Read it as missing evidence, not as a pass.
|
|
Email Triage Eval (synthetic corpus, report mode)
Email eval - Performance (TTFT / tps / latency / mem) breached its bar. The manifest ships enforce:false, so this does NOT fail the build - open the 'email-eval-report' artifact and treat it as a real regression until proven otherwise.
|
|
Email Triage Eval (synthetic corpus, report mode)
Email eval - Triage quality (FP/FN) was NOT evaluated, so this run proves nothing about it. Read it as missing evidence, not as a pass.
|
Artifacts
Produced during runtime
| Name | Size | Digest | |
|---|---|---|---|
|
email-eval-report
|
2.82 KB |
sha256:9267e3f88ca410a3eaa672c347c9d65a467a0571a7f180aad0ca8a64e045d035
|
|