You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: _posts/2026-06-09-evaluation-cards-launch.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -28,7 +28,7 @@ tags:
28
28
description: "The EvalEval Coalition launches Evaluation Cards, an open-source live interpretive layer over the AI evaluations reporting ecosystem—surfacing reproducibility, completeness, provenance, and comparability across 100,000+ reported results."
29
29
---
30
30
31
-
Today, [the EvalEval Coalition](https://evalevalai.com/) is beta launching Evaluation Cards, an open-source project for stakeholders across the evaluation ecosystem to improve reproducibility, completeness, provenance, and comparability across AI evaluations. Evaluation Cards includes:
31
+
Today, [the EvalEval Coalition](https://evalevalai.com/) is beta launching [Evaluation Cards](https://evalevalai.com/projects/eval-cards/), an open-source project for stakeholders across the evaluation ecosystem to improve reproducibility, completeness, provenance, and comparability across AI evaluations. Evaluation Cards includes:
32
32
33
33
- An interface to explore the largest and most comprehensive corpus yet assembled, with 101,955 reported evaluation results for 638 benchmarks run by 31 organizations on 5,816 models (as of 9 June 2026) based on our [EvalEval sister project, Every Eval Ever (EEE)](https://evalevalai.com/projects/every-eval-ever/)and [IBM AutoBenchmark Cards](https://research.ibm.com/publications/auto-benchmarkcard-automated-synthesis-of-benchmark-documentation).
34
34
- A front end that shows four novel signals about the state of evaluations as a whole: reproducibility, completeness, provenance, and comparability.
0 commit comments