fix(email): scoped 'anything suspicious?' query no longer dumps the full triage report #159
Triggered via pull request
August 12, 2026 17:15
Status
Failure
Total duration
2h 23m 53s
Artifacts
–
test_eval_agent_gemma_consolidation.yml
on: pull_request
rag_quality + context_retention + tool_selection vs Gemma baselines
2m 42s
Annotations
12 errors and 4 warnings
|
rag_quality + context_retention + tool_selection vs Gemma baselines
Process completed with exit code 1.
|
|
rag_quality + context_retention + tool_selection vs Gemma baselines
One or more comparisons could not be performed. Failing regardless of enforce.
|
|
rag_quality + context_retention + tool_selection vs Gemma baselines
Process completed with exit code 1.
|
|
rag_quality + context_retention + tool_selection vs Gemma baselines
Integrity: tool_selection: no scorecard produced
|
|
rag_quality + context_retention + tool_selection vs Gemma baselines
Integrity: context_retention: no scorecard produced
|
|
rag_quality + context_retention + tool_selection vs Gemma baselines
Integrity: rag_quality: no scorecard produced
|
|
rag_quality + context_retention + tool_selection vs Gemma baselines
Process completed with exit code 1.
|
|
rag_quality + context_retention + tool_selection vs Gemma baselines
No log for tool_selection - the chain stopped before it ran.
|
|
rag_quality + context_retention + tool_selection vs Gemma baselines
No log for context_retention - the chain stopped before it ran.
|
|
rag_quality + context_retention + tool_selection vs Gemma baselines
No log for rag_quality - the chain stopped before it ran.
|
|
rag_quality + context_retention + tool_selection vs Gemma baselines
Process completed with exit code 1.
|
|
rag_quality + context_retention + tool_selection vs Gemma baselines
The RAG embedder user.embeddinggemma-300m-GGUF cannot serve embeddings on this runner (probe exit 1; its output is above). rag_quality and context_retention would score ~2/10 as non-answers against a ~9/10 baseline and look like a model regression. The probe already re-registered and re-pulled the model, so a load that still fails is the RUNNER, not the registration: no Lemonade log was found on this runner (looked in: C:\windows\System32\config\systemprofile\.cache\lemonade\*.log, C:\windows\System32\config\systemprofile\.cache\lemonade\logs\*.log, C:\windows\system32\config\systemprofile\.cache\lemonade\*.log, C:\windows\system32\config\systemprofile\.cache\lemonade\logs\*.log) - read llama-server's failure from the Lemonade console instead, then on sjlab-stx-halo-18 run `powershell -File installer\scripts\ensure-lemonade-running.ps1 -ForceRestart` and re-run this probe by hand (`python tests/ci_lemonade_check.py --model user.embeddinggemma-300m-GGUF --checkpoint ggml-org/embeddinggemma-300M-GGUF:Q8_0 --recipe llamacpp --register-embedding --embeddings --llamacpp-args '--ubatch-size 2048'`). Do NOT re-run the eval until it passes - it costs 3.5h and measures nothing.
|
|
rag_quality + context_retention + tool_selection vs Gemma baselines
No files were found with the provided path: eval-out/. No artifacts will be uploaded.
|
|
rag_quality + context_retention + tool_selection vs Gemma baselines
tool_selection: no scorecard produced - nothing to compare.
|
|
rag_quality + context_retention + tool_selection vs Gemma baselines
context_retention: no scorecard produced - nothing to compare.
|
|
rag_quality + context_retention + tool_selection vs Gemma baselines
rag_quality: no scorecard produced - nothing to compare.
|