braintrustdata ★ 986 Repository stable Open source tool for evaluating AI model outputs using established best practices.
Best match because its name, category, capabilities, or owner matches “evals”. Adoption and trust break close ties.
Stars 986
Downloads —
Updated 2026-07-29
Skill installs — Trust pending Momentum not matched
Evals
openai ★ 19k Repository experimental Framework for creating and running evaluations on large language models and systems.
Best match because its name, category, capabilities, or owner matches “evals”. Adoption and trust break close ties.
Stars 19k
Downloads 2.5k
Updated 2026-04-14
Skill installs — Trust pending Momentum not matched
$ pip install evalsCopy Evals
Open-source LLM evaluation, tracing, and monitoring by Comet.
Best match because its name, category, capabilities, or owner matches “evals”. Adoption and trust break close ties.
Stars 21k
Downloads 3527k
Updated 2026-08-03
Skill installs — Trust pending Momentum not matched
$ pip install opikCopy Tracing & Observability Evals
Open-source LLM tracing, evaluation, and observability platform.
Best match because its name, category, capabilities, or owner matches “evals”. Adoption and trust break close ties.
Stars 11k
Downloads 2215k
Updated 2026-08-03
Skill installs — Trust pending Momentum not matched
$ pip install arize-phoenixCopy Tracing & Observability
wandb ★ 1.1k Tool experimental Toolkit to track, evaluate, and monitor LLM applications by W&B.
Best match because its name, category, capabilities, or owner matches “evals”. Adoption and trust break close ties.
Stars 1.1k
Downloads 947k
Updated 2026-08-01
Skill installs — Trust pending Momentum not matched
$ pip install weaveCopy Experiment Tracking
promptfoo ★ 24k Tool experimental Test, evaluate, and red-team LLM prompts and apps from the CLI.
Best match because its name, category, capabilities, or owner matches “evals”. Adoption and trust break close ties.
Stars 24k
Downloads —
Updated 2026-08-03
Skill installs — Trust pending Momentum not matched
$ npm i promptfooCopy Evals
evidentlyai ★ 7.8k Tool experimental Open source framework to evaluate and monitor machine learning and LLM systems.
Best match because its name, category, capabilities, or owner matches “evals”. Adoption and trust break close ties.
Stars 7.8k
Downloads —
Updated 2026-05-02
Skill installs — Trust pending Momentum not matched
Evals
confident-ai ★ 17k Tool stable Open source evaluation framework for unit testing large language model outputs.
Best match because its name, category, capabilities, or owner matches “evals”. Adoption and trust break close ties.
Stars 17k
Downloads 6221k
Updated 2026-08-02
Skill installs — Trust pending Momentum not matched
$ pip install deepevalCopy Evals
EleutherAI ★ 14k Repository experimental Unified framework for evaluating language models across hundreds of benchmark tasks.
Best match because its name, category, capabilities, or owner matches “evals”. Adoption and trust break close ties.
Stars 14k
Downloads 1512k
Updated 2026-07-13
Skill installs — Trust pending Momentum not matched
$ pip install lm_evalCopy Evals
truera ★ 3.5k Repository stable Library for evaluating and tracking the performance of LLM applications with feedback functions.
Best match because its name, category, capabilities, or owner matches “evals”. Adoption and trust break close ties.
Stars 3.5k
Downloads 97k
Updated 2026-07-31
Skill installs — Trust pending Momentum not matched
$ pip install trulensCopy Evals Tracing & Observability
Toolkit for building, evaluating and deploying prompt based LLM application workflows.
Best match because its name, category, capabilities, or owner matches “evals”. Adoption and trust break close ties.
Stars 11k
Downloads 79k
Updated 2026-07-09
Skill installs — Trust pending Momentum not matched
$ pip install promptflowCopy Prompt Management Evals
Giskard-AI ★ 5.7k Repository stable Open source testing framework to detect vulnerabilities and evaluate quality of LLM and ML models.
Best match because its name, category, capabilities, or owner matches “evals”. Adoption and trust break close ties.
Stars 5.7k
Downloads 25k
Updated 2026-08-03
Skill installs — Trust pending Momentum not matched
$ pip install giskardCopy Evals
explodinggradients ★ 15k Tool experimental Framework for evaluating retrieval augmented generation pipelines with reference free metrics.
Best match because its name, category, capabilities, or owner matches “evals”. Adoption and trust break close ties.
Stars 15k
Downloads 1582k
Updated 2026-02-24
Skill installs — Trust pending Momentum not matched
$ pip install ragasCopy Evals
hegelai ★ 3.0k Tool experimental Open source tools for testing and experimenting with LLM prompts and vector databases.
Best match because its name, category, capabilities, or owner matches “evals”. Adoption and trust break close ties.
Stars 3.0k
Downloads 82
Updated 2026-02-11
Skill installs — Trust pending Momentum not matched
$ pip install prompttoolsCopy Prompt Management Evals