Control center / Discover

Discover

Browse by category instead of drowning in search results. Every category opens a submenu of subcategories with live counts, and checking one narrows the grid to exactly that slice.

Verified against GitHub · scores pending
4 results for “benchmark” across 3 categories
Sort by:Type:Trust:Maturity:4 shown

Security

1

kube-bench

aquasecurity8.1kToolexperimental

Checks Kubernetes clusters against the CIS Benchmark.

Best match because its name, category, capabilities, or owner matches “benchmark”. Adoption and trust break close ties.

Stars
8.1k
Downloads
Updated
2026-07-30
Skill installs
Trust pendingMomentum not matched

AI Agents

1

SWE-agent

princeton-nlp20kAgentactive

Autonomous software engineering agent designed to solve real GitHub issues in SWE-bench.

Best match because its name, category, capabilities, or owner matches “benchmark”. Adoption and trust break close ties.

Stars
20k
Downloads
Updated
2026-07-27
Skill installs
Trust pendingMomentum not matched

LLM Ops & Observability

2

LM Evaluation Harness

EleutherAI14kRepositoryexperimental

Unified framework for evaluating language models across hundreds of benchmark tasks.

Best match because its name, category, capabilities, or owner matches “benchmark”. Adoption and trust break close ties.

Stars
14k
Downloads
1512k
Updated
2026-07-13
Skill installs
Trust pendingMomentum not matched

OpenAI Evals

openai19kRepositoryexperimental

Framework for creating and running evaluations on large language models and systems.

Best match because its name, category, capabilities, or owner matches “benchmark”. Adoption and trust break close ties.

Stars
19k
Downloads
2.5k
Updated
2026-04-14
Skill installs
Trust pendingMomentum not matched
Discover · SkillPilot