Control center / Discover

Discover

Browse by category instead of drowning in search results. Every category opens a submenu of subcategories with live counts, and checking one narrows the grid to exactly that slice.

Verified against GitHub · scores pending
21 results for “inference” across 3 categories
Sort by:Type:Trust:Maturity:21 shown

Token & Cost Optimization

14

Text Generation Inference

huggingface★ 11kToolslowing

Production ready inference server for LLMs optimized for throughput and cost efficiency.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
11k
Downloads
—
Updated
2026-03-21
Skill installs
—
Trust pendingMomentum not matched

SGLang

sgl-project★ 36kRepositoryexperimental

Fast serving framework for LLMs with radix attention based prefix caching for lower cost.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
36k
Downloads
100357k
Updated
2026-09-28
Skill installs
—
Trust pendingMomentum not matched

vLLM

vllm-project★ 93kRepositoryexperimental

High throughput and memory efficient inference and serving engine for large language models.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
93k
Downloads
5581k
Updated
2026-09-28
Skill installs
—
Trust pendingMomentum not matched

Hugging Face Optimum

huggingface★ 3.5kRepositorystable

Toolkit for optimizing transformer models for faster and cheaper inference on various hardware.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
3.5k
Downloads
1991k
Updated
2026-09-24
Skill installs
—
Trust pendingMomentum not matched

llama-cpp-python

abetlen★ 11kRepositoryexperimental

Python bindings for llama.cpp enabling low cost local large language model inference.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
11k
Downloads
698k
Updated
2026-09-22
Skill installs
—
Trust pendingMomentum not matched

Ollama

ollama★ 182kToolexperimental

Run large language models locally to reduce inference costs and latency.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
182k
Downloads
—
Updated
2026-09-27
Skill installs
—
Trust pendingMomentum not matched

llama.cpp

ggerganov★ 130kRepositorystable

Efficient C plus plus inference of LLaMA models enabling cheaper on device execution.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
130k
Downloads
—
Updated
2026-09-27
Skill installs
—
Trust pendingMomentum not matched

LMCache

LMCache★ 12kToolstable

KV cache management layer that speeds up LLM serving by reusing cached context.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
12k
Downloads
68k
Updated
2026-09-28
Skill installs
—
Trust pendingMomentum not matched

DeepSpeed

microsoft★ 43kRepositoryexperimental

Deep learning optimization library enabling efficient large model training and inference.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
43k
Downloads
—
Updated
2026-09-28
Skill installs
—
Trust pendingMomentum not matched

llamafile

Mozilla-Ocho★ 26kToolexperimental

Single file executables that run LLMs locally without heavy infrastructure or API cost.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
26k
Downloads
—
Updated
2026-09-25
Skill installs
—
Trust pendingMomentum not matched

TensorRT-LLM

NVIDIA★ 15kRepositorystable

Toolkit for optimizing and deploying large language models with high performance inference.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
15k
Downloads
12k
Updated
2026-09-28
Skill installs
—
Trust pendingMomentum not matched

OptiLLM

codelion★ 4.3kToolexperimental

Optimizing inference proxy that applies techniques to improve accuracy and reduce LLM cost.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
4.3k
Downloads
—
Updated
2026-09-28
Skill installs
—
Trust pendingMomentum not matched

ExLlamaV2

turboderp★ 4.6kRepositoryslowing

Fast inference library for running quantized LLMs on consumer GPUs at lower cost.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
4.6k
Downloads
—
Updated
2026-03-04
Skill installs
—
Trust pendingMomentum not matched

TensorZero

tensorzero★ 12kToolactive

Open source LLM gateway and optimization framework unifying inference, observability and evals.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
12k
Downloads
—
Updated
2026-06-11
Skill installs
—
Trust pendingMomentum not matched

LLM Ops & Observability

6

OpenInference

Arize-ai★ 1.2kRepositorystable

Open standard for instrumenting LLM applications to capture traces and spans for observability.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
1.2k
Downloads
1559k
Updated
2026-09-28
Skill installs
—
Trust pendingMomentum not matched

LocalAI

mudler★ 49kToolstable

A drop in replacement REST API that is compatible with OpenAI for local LLM inferencing.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
49k
Downloads
—
Updated
2026-09-28
Skill installs
—
Trust pendingMomentum not matched

Exo

exo-explore★ 48kToolactive

Run your own AI cluster at home with everyday consumer devices to serve models locally.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
48k
Downloads
—
Updated
2026-09-28
Skill installs
—
Trust pendingMomentum not matched

Outlines

dottxt-ai★ 16kToolstable

Structured text generation library for large language models.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
16k
Downloads
—
Updated
2026-09-21
Skill installs
—
Trust pendingMomentum not matched

LMDeploy

InternLM★ 8.1kToolexperimental

A toolkit for compressing deploying and serving large language models locally with high throughput.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
8.1k
Downloads
—
Updated
2026-09-23
Skill installs
—
Trust pendingMomentum not matched

Mistral.rs

EricLBuehler★ 7.7kToolexperimental

A fast LLM inference engine written in Rust for local model deployment.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
7.7k
Downloads
—
Updated
2026-09-25
Skill installs
—
Trust pendingMomentum not matched

RAG & Knowledge

1

Text Embeddings Inference

huggingface★ 5.1kToolstable

Fast inference server for serving text embedding models in production.

Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.

Stars
5.1k
Downloads
—
Updated
2026-09-23
Skill installs
—
Trust pendingMomentum not matched
Discover · SkillPilot