Popular repositories Loading
-
every_eval_ever
every_eval_ever PublicForked from evaleval/every_eval_ever
Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and research papers to loca…
Python
-
trl
trl PublicForked from huggingface/trl
Train transformer language models with reinforcement learning.
Python
-
gorilla
gorilla PublicForked from ShishirPatil/gorilla
Gorilla: Training and Evaluating LLMs for Function Calls (Tool Calls)
Python
-
evals
evals PublicForked from openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Python
-
BIG-bench
BIG-bench PublicForked from google/BIG-bench
Beyond the Imitation Game collaborative benchmark for measuring and extrapolating the capabilities of language models
Python
-
lm-evaluation-harness
lm-evaluation-harness PublicForked from EleutherAI/lm-evaluation-harness
A framework for few-shot evaluation of language models.
Python
If the problem persists, check the GitHub status page or contact support.
