Control center / Discover

Discover

Browse by category instead of drowning in search results. Every category opens a submenu of subcategories with live counts, and checking one narrows the grid to exactly that slice.

Verified against GitHub · scores pending
11 results for “evaluation” across 2 categories
Sort by:Type:Trust:Maturity:11 shown

LLM Ops & Observability

10

LM Evaluation Harness

EleutherAI14kRepositoryexperimental

Unified framework for evaluating language models across hundreds of benchmark tasks.

Best match because its name, category, capabilities, or owner matches “evaluation”. Adoption and trust break close ties.

Stars
14k
Downloads
1512k
Updated
2026-07-13
Skill installs
Trust pendingMomentum not matched

DeepEval

confident-ai17kToolstable

Open source evaluation framework for unit testing large language model outputs.

Best match because its name, category, capabilities, or owner matches “evaluation”. Adoption and trust break close ties.

Stars
17k
Downloads
6221k
Updated
2026-08-02
Skill installs
Trust pendingMomentum not matched

TruLens

truera3.5kRepositorystable

Library for evaluating and tracking the performance of LLM applications with feedback functions.

Best match because its name, category, capabilities, or owner matches “evaluation”. Adoption and trust break close ties.

Stars
3.5k
Downloads
97k
Updated
2026-07-31
Skill installs
Trust pendingMomentum not matched

Giskard

Giskard-AI5.7kRepositorystable

Open source testing framework to detect vulnerabilities and evaluate quality of LLM and ML models.

Best match because its name, category, capabilities, or owner matches “evaluation”. Adoption and trust break close ties.

Stars
5.7k
Downloads
25k
Updated
2026-08-03
Skill installs
Trust pendingMomentum not matched

Agenta

agenta-ai4.4kToolexperimental

Open-source LLMOps platform for prompt management and evaluation.

Best match because its name, category, capabilities, or owner matches “evaluation”. Adoption and trust break close ties.

Stars
4.4k
Downloads
16k
Updated
2026-08-03
Skill installs
Trust pendingMomentum not matched

Ragas

explodinggradients15kToolexperimental

Framework for evaluating retrieval augmented generation pipelines with reference free metrics.

Best match because its name, category, capabilities, or owner matches “evaluation”. Adoption and trust break close ties.

Stars
15k
Downloads
1582k
Updated
2026-02-24
Skill installs
Trust pendingMomentum not matched

OpenAI Evals

openai19kRepositoryexperimental

Framework for creating and running evaluations on large language models and systems.

Best match because its name, category, capabilities, or owner matches “evaluation”. Adoption and trust break close ties.

Stars
19k
Downloads
2.5k
Updated
2026-04-14
Skill installs
Trust pendingMomentum not matched

PromptTools

hegelai3.0kToolexperimental

Open source tools for testing and experimenting with LLM prompts and vector databases.

Best match because its name, category, capabilities, or owner matches “evaluation”. Adoption and trust break close ties.

Stars
3.0k
Downloads
82
Updated
2026-02-11
Skill installs
Trust pendingMomentum not matched

Opik

comet-ml21kToolstable

Open-source LLM evaluation, tracing, and monitoring by Comet.

Best match because its name, category, capabilities, or owner matches “evaluation”. Adoption and trust break close ties.

Stars
21k
Downloads
3527k
Updated
2026-08-03
Skill installs
Trust pendingMomentum not matched

Phoenix

Arize-ai11kToolstable

Open-source LLM tracing, evaluation, and observability platform.

Best match because its name, category, capabilities, or owner matches “evaluation”. Adoption and trust break close ties.

Stars
11k
Downloads
2215k
Updated
2026-08-03
Skill installs
Trust pendingMomentum not matched

RAG & Knowledge

1

AutoRAG

Marker-Inc-Korea5.0kRepositorystable

Automated tool for evaluating and optimizing RAG pipeline configurations.

Best match because its name, category, capabilities, or owner matches “evaluation”. Adoption and trust break close ties.

Stars
5.0k
Downloads
300
Updated
2026-08-01
Skill installs
Trust pendingMomentum not matched
Discover · SkillPilot