Control center / Discover

Discover

Browse by category instead of drowning in search results. Every category opens a submenu of subcategories with live counts, and checking one narrows the grid to exactly that slice.

Verified against GitHub · scores pending
1 result for “benchmarking” across 1 categories
Sort by:Type:Trust:Maturity:1 shown

LLM Ops & Observability

1

OpenAI Evals

openai19kRepositoryexperimental

Framework for creating and running evaluations on large language models and systems.

Best match because its name, category, capabilities, or owner matches “benchmarking”. Adoption and trust break close ties.

Stars
19k
Downloads
2.5k
Updated
2026-04-14
Skill installs
Trust pendingMomentum not matched
Discover · SkillPilot