huggingface ★ 11k Tool active Production ready inference server for LLMs optimized for throughput and cost efficiency.
Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.
Stars 11k
Downloads —
Updated 2026-03-21
Skill installs — Trust pending Momentum not matched
Observability & Cost
sgl-project ★ 31k Repository experimental Fast serving framework for LLMs with radix attention based prefix caching for lower cost.
Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.
Stars 31k
Downloads 100357k
Updated 2026-08-03
Skill installs — Trust pending Momentum not matched
$ pip install sglangCopy Prompt Caching
vllm-project ★ 88k Repository experimental High throughput and memory efficient inference and serving engine for large language models.
Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.
Stars 88k
Downloads 5581k
Updated 2026-08-03
Skill installs — Trust pending Momentum not matched
$ pip install vllmCopy Model Routing Observability & Cost
huggingface ★ 3.5k Repository stable Toolkit for optimizing transformer models for faster and cheaper inference on various hardware.
Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.
Stars 3.5k
Downloads 1991k
Updated 2026-07-31
Skill installs — Trust pending Momentum not matched
$ pip install optimumCopy Context Compression
abetlen ★ 11k Repository experimental Python bindings for llama.cpp enabling low cost local large language model inference.
Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.
Stars 11k
Downloads 698k
Updated 2026-08-02
Skill installs — Trust pending Momentum not matched
$ pip install llama_cpp_pythonCopy Cheaper-Model Agents
ollama ★ 178k Tool experimental Run large language models locally to reduce inference costs and latency.
Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.
Stars 178k
Downloads —
Updated 2026-07-31
Skill installs — Trust pending Momentum not matched
Cheaper-Model Agents
ggerganov ★ 122k Repository stable Efficient C plus plus inference of LLaMA models enabling cheaper on device execution.
Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.
Stars 122k
Downloads —
Updated 2026-08-02
Skill installs — Trust pending Momentum not matched
$ pip install llama-cpp-scriptsCopy Cheaper-Model Agents
KV cache management layer that speeds up LLM serving by reusing cached context.
Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.
Stars 11k
Downloads 68k
Updated 2026-08-02
Skill installs — Trust pending Momentum not matched
$ pip install lmcacheCopy Prompt Caching
microsoft ★ 43k Repository experimental Deep learning optimization library enabling efficient large model training and inference.
Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.
Stars 43k
Downloads —
Updated 2026-08-03
Skill installs — Trust pending Momentum not matched
Context Compression
Mozilla-Ocho ★ 25k Tool experimental Single file executables that run LLMs locally without heavy infrastructure or API cost.
Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.
Stars 25k
Downloads —
Updated 2026-07-31
Skill installs — Trust pending Momentum not matched
Cheaper-Model Agents
NVIDIA ★ 14k Repository stable Toolkit for optimizing and deploying large language models with high performance inference.
Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.
Stars 14k
Downloads 12k
Updated 2026-08-03
Skill installs — Trust pending Momentum not matched
$ pip install tensorrt_llmCopy Model Routing
codelion ★ 4.2k Tool experimental Optimizing inference proxy that applies techniques to improve accuracy and reduce LLM cost.
Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.
Stars 4.2k
Downloads —
Updated 2026-07-18
Skill installs — Trust pending Momentum not matched
$ pip install optillmCopy LLM Gateways
turboderp ★ 4.6k Repository experimental Fast inference library for running quantized LLMs on consumer GPUs at lower cost.
Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.
Stars 4.6k
Downloads —
Updated 2026-03-04
Skill installs — Trust pending Momentum not matched
$ pip install exllamav2_extCopy Cheaper-Model Agents
tensorzero ★ 12k Tool stable Open source LLM gateway and optimization framework unifying inference, observability and evals.
Best match because its name, category, capabilities, or owner matches “inference”. Adoption and trust break close ties.
Stars 12k
Downloads —
Updated 2026-06-11
Skill installs — Trust pending Momentum not matched
LLM Gateways Observability & Cost