Control center / Discover

Discover

Browse by category instead of drowning in search results. Every category opens a submenu of subcategories with live counts, and checking one narrows the grid to exactly that slice.

Verified against GitHub · scores pending
15 results for “Evals” across 2 categories
Sort by:Type:Trust:Maturity:15 shown

Token & Cost Optimization

1

TensorZero

tensorzero12kToolstable

Open source LLM gateway and optimization framework unifying inference, observability and evals.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
12k
Downloads
Updated
2026-06-11
Skill installs
Trust pendingMomentum not matched

LLM Ops & Observability

14

AutoEvals

braintrustdata986Repositorystable

Open source tool for evaluating AI model outputs using established best practices.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
986
Downloads
Updated
2026-07-29
Skill installs
Trust pendingMomentum not matched

OpenAI Evals

openai19kRepositoryexperimental

Framework for creating and running evaluations on large language models and systems.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
19k
Downloads
2.5k
Updated
2026-04-14
Skill installs
Trust pendingMomentum not matched

Opik

comet-ml21kToolstable

Open-source LLM evaluation, tracing, and monitoring by Comet.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
21k
Downloads
3527k
Updated
2026-08-03
Skill installs
Trust pendingMomentum not matched

Phoenix

Arize-ai11kToolstable

Open-source LLM tracing, evaluation, and observability platform.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
11k
Downloads
2215k
Updated
2026-08-03
Skill installs
Trust pendingMomentum not matched

Weave

wandb1.1kToolexperimental

Toolkit to track, evaluate, and monitor LLM applications by W&B.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
1.1k
Downloads
947k
Updated
2026-08-01
Skill installs
Trust pendingMomentum not matched

Promptfoo

promptfoo24kToolexperimental

Test, evaluate, and red-team LLM prompts and apps from the CLI.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
24k
Downloads
Updated
2026-08-03
Skill installs
Trust pendingMomentum not matched

Evidently

evidentlyai7.8kToolexperimental

Open source framework to evaluate and monitor machine learning and LLM systems.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
7.8k
Downloads
Updated
2026-05-02
Skill installs
Trust pendingMomentum not matched

DeepEval

confident-ai17kToolstable

Open source evaluation framework for unit testing large language model outputs.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
17k
Downloads
6221k
Updated
2026-08-02
Skill installs
Trust pendingMomentum not matched

LM Evaluation Harness

EleutherAI14kRepositoryexperimental

Unified framework for evaluating language models across hundreds of benchmark tasks.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
14k
Downloads
1512k
Updated
2026-07-13
Skill installs
Trust pendingMomentum not matched

TruLens

truera3.5kRepositorystable

Library for evaluating and tracking the performance of LLM applications with feedback functions.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
3.5k
Downloads
97k
Updated
2026-07-31
Skill installs
Trust pendingMomentum not matched

Prompt Flow

microsoft11kToolactive

Toolkit for building, evaluating and deploying prompt based LLM application workflows.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
11k
Downloads
79k
Updated
2026-07-09
Skill installs
Trust pendingMomentum not matched

Giskard

Giskard-AI5.7kRepositorystable

Open source testing framework to detect vulnerabilities and evaluate quality of LLM and ML models.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
5.7k
Downloads
25k
Updated
2026-08-03
Skill installs
Trust pendingMomentum not matched

Ragas

explodinggradients15kToolexperimental

Framework for evaluating retrieval augmented generation pipelines with reference free metrics.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
15k
Downloads
1582k
Updated
2026-02-24
Skill installs
Trust pendingMomentum not matched

PromptTools

hegelai3.0kToolexperimental

Open source tools for testing and experimenting with LLM prompts and vector databases.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
3.0k
Downloads
82
Updated
2026-02-11
Skill installs
Trust pendingMomentum not matched
Discover · SkillPilot