microsoft ★ 6.5k Tool experimental Prompt compression technique to shrink context length by up to 20x with minimal performance loss.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 6.5k
Downloads 93k
Updated 2026-04-08
Skill installs — Trust pending Momentum not matched
$ pip install llmlinguaCopy Context Compression
TensorOpsAI ★ 388 Tool experimental Gateway and management tool for calling, caching and monitoring multiple LLM providers.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 388
Downloads —
Updated 2026-07-29
Skill installs — Trust pending Momentum not matched
$ pip install llmstudio-monorepoCopy LLM Gateways
Call 100+ LLM APIs using OpenAI format with proxy cost tracking and budget caps.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 55k
Downloads 661278k
Updated 2026-08-03
Skill installs — Trust pending Momentum not matched
$ pip install litellmCopy LLM Gateways
vllm-project ★ 88k Repository experimental High throughput and memory efficient inference and serving engine for large language models.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 88k
Downloads 5581k
Updated 2026-08-03
Skill installs — Trust pending Momentum not matched
$ pip install vllmCopy Model Routing Observability & Cost
NVIDIA ★ 14k Repository stable Toolkit for optimizing and deploying large language models with high performance inference.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 14k
Downloads 12k
Updated 2026-08-03
Skill installs — Trust pending Momentum not matched
$ pip install tensorrt_llmCopy Model Routing
codelion ★ 4.2k Tool experimental Optimizing inference proxy that applies techniques to improve accuracy and reduce LLM cost.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 4.2k
Downloads —
Updated 2026-07-18
Skill installs — Trust pending Momentum not matched
$ pip install optillmCopy LLM Gateways
bentoml ★ 12k Tool experimental Platform for running and deploying open source LLMs cost effectively in production.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 12k
Downloads 1.8k
Updated 2026-07-27
Skill installs — Trust pending Momentum not matched
$ pip install openllmCopy Cheaper-Model Agents
Framework for serving and routing queries between strong and weak LLMs to slash costs by 85%.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 5.3k
Downloads 9.3k
Updated 2024-08-10
Skill installs — Trust pending Momentum not matched
$ pip install routellmCopy Model Routing
abetlen ★ 11k Repository experimental Python bindings for llama.cpp enabling low cost local large language model inference.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 11k
Downloads 698k
Updated 2026-08-02
Skill installs — Trust pending Momentum not matched
$ pip install llama_cpp_pythonCopy Cheaper-Model Agents
ollama ★ 178k Tool experimental Run large language models locally to reduce inference costs and latency.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 178k
Downloads —
Updated 2026-07-31
Skill installs — Trust pending Momentum not matched
Cheaper-Model Agents
ggerganov ★ 122k Repository stable Efficient C plus plus inference of LLaMA models enabling cheaper on device execution.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 122k
Downloads —
Updated 2026-08-02
Skill installs — Trust pending Momentum not matched
$ pip install llama-cpp-scriptsCopy Cheaper-Model Agents
Mozilla-Ocho ★ 25k Tool experimental Single file executables that run LLMs locally without heavy infrastructure or API cost.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 25k
Downloads —
Updated 2026-07-31
Skill installs — Trust pending Momentum not matched
Cheaper-Model Agents
tensorzero ★ 12k Tool stable Open source LLM gateway and optimization framework unifying inference, observability and evals.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 12k
Downloads —
Updated 2026-06-11
Skill installs — Trust pending Momentum not matched
LLM Gateways Observability & Cost
token-js ★ 310 Tool experimental Unified TypeScript SDK for calling over a hundred LLMs through one API.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 310
Downloads 3.7k
Updated 2026-07-05
Skill installs — Trust pending Momentum not matched
$ npm i token.jsCopy LLM Gateways
langwatch ★ 3.5k Tool stable Monitoring and analytics platform for tracking LLM usage quality and cost.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 3.5k
Downloads —
Updated 2026-08-03
Skill installs — Trust pending Momentum not matched
$ npm i @langwatch/serverCopy Observability & Cost
lm-sys ★ 40k Repository experimental Platform for training, serving and evaluating multiple open LLMs with routing capabilities.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 40k
Downloads 45k
Updated 2026-05-01
Skill installs — Trust pending Momentum not matched
$ pip install fschatCopy Model Routing
Unified LLM gateway with provider fallbacks, prompt routing, and spend observability.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 26k
Downloads —
Updated 2026-08-02
Skill installs 45k Trust pending Momentum not matched
LLM Gateways
Portkey-AI ★ 13k Tool stable Fast open-source AI gateway with automated retries, semantic caching, and budget rules.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 13k
Downloads 4.1k
Updated 2026-05-25
Skill installs — Trust pending Momentum not matched
$ npm i @portkey-ai/gatewayCopy LLM Gateways
openrouter-ai Tool experimental Unified LLM gateway with dynamic price routing, prompt caching, and cost analytics.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars —
Downloads —
Updated —
Skill installs — Trust pending Momentum not matched
LLM Gateways Model Routing
sgl-project ★ 31k Repository experimental Fast serving framework for LLMs with radix attention based prefix caching for lower cost.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 31k
Downloads 100357k
Updated 2026-08-03
Skill installs — Trust pending Momentum not matched
$ pip install sglangCopy Prompt Caching
Open-source LLM engineering platform for tracing, prompt management, and per-user cost tracking.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 32k
Downloads 7102k
Updated 2026-08-01
Skill installs — Trust pending Momentum not matched
$ npm i langfuseCopy Observability & Cost
skypilot-org ★ 10k Tool experimental Framework for running LLM workloads across clouds to minimize compute cost and maximize availability.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 10k
Downloads 1705k
Updated 2026-08-03
Skill installs — Trust pending Momentum not matched
$ pip install skypilotCopy Observability & Cost
aurelio-labs ★ 3.8k Tool experimental Superfast decision layer for routing queries to the right LLM or logic using semantic similarity.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 3.8k
Downloads 441k
Updated 2026-07-26
Skill installs — Trust pending Momentum not matched
$ pip install semantic-routerCopy Model Routing
openmeterio ★ 2.2k Tool stable Usage metering and billing infrastructure for tracking LLM API consumption and cost.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 2.2k
Downloads 74k
Updated 2026-08-02
Skill installs — Trust pending Momentum not matched
$ pip install openmeterCopy Observability & Cost
KV cache management layer that speeds up LLM serving by reusing cached context.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 11k
Downloads 68k
Updated 2026-08-02
Skill installs — Trust pending Momentum not matched
$ pip install lmcacheCopy Prompt Caching
zilliztech ★ 8.1k Tool slowing Semantic cache for storing LLM responses to bypass redundant API calls.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 8.1k
Downloads 437k
Updated 2025-07-11
Skill installs — Trust pending Momentum not matched
$ pip install gptcacheCopy Prompt Caching
Open-source LLM observability platform with real-time token tracking, caching, and cost analytics.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 6.0k
Downloads 3.9k
Updated 2026-07-25
Skill installs — Trust pending Momentum not matched
$ npm i @helicone/heliconeCopy Observability & Cost
predibase ★ 3.8k Repository active Framework for serving thousands of fine tuned LLM adapters on shared GPU infrastructure.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 3.8k
Downloads —
Updated 2026-05-28
Skill installs — Trust pending Momentum not matched
Model Routing
huggingface ★ 11k Tool active Production ready inference server for LLMs optimized for throughput and cost efficiency.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 11k
Downloads —
Updated 2026-03-21
Skill installs — Trust pending Momentum not matched
Observability & Cost
turboderp ★ 4.6k Repository experimental Fast inference library for running quantized LLMs on consumer GPUs at lower cost.
Best match because its name, category, capabilities, or owner matches “llm”. Adoption and trust break close ties.
Stars 4.6k
Downloads —
Updated 2026-03-04
Skill installs — Trust pending Momentum not matched
$ pip install exllamav2_extCopy Cheaper-Model Agents