Discover / Token & Cost Optimization

LoRAX

by predibasePython

Framework for serving thousands of fine tuned LLM adapters on shared GPU infrastructure.

Repositoryactive

Maturity: active because commit 67d ago, latest release lorax-0.4.0. Derived from release and commit history, not a rating.

Stars
3.8k
Forks
326
Downloads / mo
Last commit
2026-05-28
License
Apache-2.0
Open issues
184

Market and trust evidence

Edition not yet matched

No exact skills.sh identity match is available for this repository. Repository adoption and freshness remain visible above; install momentum is not inferred.

Trust analysis is a screening signal, not a security warranty. Read the ranking and trust methodology.

In practice

Written by AI from this repository’s README · high confidence

Running one server per fine tuned model multiplies GPU cost even though the adapters share a base model.

Use it when

Use it when many LoRA adapters share one base model and you want them loaded per request on shared hardware.

Not the right pick when

It requires an Ampere or newer Nvidia GPU on Linux with Docker, so other hardware is out of scope.

Capabilities

  • dynamic adapter loading just in time per request
  • heterogeneous continuous batching across adapters
  • adapter exchange scheduling between GPU and CPU memory
  • tensor parallelism and prebuilt CUDA kernels
  • OpenAI compatible API with multi turn chat and structured output
  • Helm charts, Prometheus metrics and OpenTelemetry tracing

Requirements

  • Nvidia GPU (Ampere generation or above)
  • CUDA 11.8 compatible device drivers and above
  • Linux OS and Docker with nvidia-container-toolkit

Cost: Free and open source

Video walkthroughs

Third-party YouTube uploads matched to this tool by title, channel and repository name on 2026-08-03. Not made, reviewed or endorsed by SkillPilot. View counts and publish months are as of the match date and the month is approximate. Nothing loads from YouTube until you press play.

What the repository ships

Has testsHas docsDocker imageCI configured

Detected from the actual files in the repository root.

Latest release lorax-0.4.0

Published 2025-01-13

LoRAX is the open-source framework for serving hundreds of fine-tuned LLMs in production for the price of one.

Tags

README

<p align="center">

<a href="https://github.com/predibase/lorax">

<img src="docs/LoRAX_Main_Logo-Orange.png" alt="LoRAX Logo" style="width:200px;" />

</a>

</p>

<div align="center">

_LoRAX: Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs_

image

License

Artifact Hub

</div>

LoRAX (LoRA eXchange) is a framework that allows users to serve thousands of fine-tuned models on a single GPU, dramatically reducing the cost of serving without compromising on throughput or latency.

📖 Table of contents

  • 📖 Table of contents
  • 🌳 Features
  • 🏠 Models
  • 🏃‍♂️ Getting Started
  • Requirements
  • Launch LoRAX Server
  • Prompt via REST API
  • Prompt via Python Client
  • Chat via OpenAI API
  • Next steps
  • 🙇 Acknowledgements
  • 🗺️ Roadmap

🌳 Features

  • 🚅 Dynamic Adapter Loading: include any fine-tuned LoRA adapter from HuggingFace, Predibase, or any filesystem in your request, it will be loaded just-in-time without blocking concurrent requests. Merge adapters per request to instantly create powerful ensembles.
  • 🏋️‍♀️ Heterogeneous Continuous Batching: packs requests for different adapters together into the same batch, keeping latency and throughput nearly constant with the number of concurrent adapters.
  • 🧁 Adapter Exchange Scheduling: asynchronously prefetches and offloads adapters between GPU and CPU memory, schedules request batching to optimize the aggregate throughput of the system.
  • 👬 Optimized Inference: high throughput and low latency optimizations including tensor parallelism, pre-compiled CUDA kernels (flash-attention, paged attention, SGMV), quantization, token streaming.
  • 🚢 Ready for Production prebuilt Docker images, Helm charts for Kubernetes, Prometheus metrics, and distributed tracing with Open Telemetry. OpenAI compatible API supporting multi-turn chat conversations. Private adapters through per-request tenant isolation. Structured Output (JSON mode).
  • 🤯 Free for Commercial Use: Apache 2.0 License. Enough said 😎.

<p align="center">

<img src="https://github.com/predibase/lorax/assets/29719151/f88aa16c-66de-45ad-ad40-01a7874ed8a9" />

</p>

🏠 Models

Serving a fine-tuned model with LoRAX consists of two components:

  • Base Model: pretrained large model shared across all adapters.
  • Adapter: task-specific adapter weights dynamically loaded per request.

LoRAX supports a number of Large Language Models as the base model including Llama (including CodeLlama), Mistral (including Zephyr), and Qwen. See Supported Architectures for a complete list of supported base models.

Base models can be loaded in fp16 or quantized with bitsandbytes, GPT-Q, or AWQ.

Supported adapters include LoRA adapters trained using the PEFT and Ludwig libraries. Any of the linear layers in the model can be adapted via LoRA and loaded in LoRAX.

🏃‍♂️ Getting Started

We recommend starting with our pre-built Docker image to avoid compiling custom CUDA kernels and other dependencies.

Requirements

The minimum system requirements need to run LoRAX include:

  • Nvidia GPU (Ampere generation or above)
  • CUDA 11.8 compatible device drivers and above
  • Linux OS
  • Docker (for this guide)

Launch LoRAX Server

Prerequisites

Install nvidia-container-toolkit

Then

  • sudo systemctl daemon-reload
  • sudo systemctl restart docker

model=mistralai/Mistral-7B-Instruct-v0.1
volume=$PWD/data

docker run --gpus all --shm-size 1g -p 8080:80 -v $volume:/data \
    ghcr.io/predibase/lorax:main --model-id $model

For a full tutorial including token streaming and the Python client, see Getting Started - Docker.

Prompt via REST API

Prompt base LLM:


curl 127.0.0.1:8080/generate \
    -X POST \
    -d '{
        "inputs": "[INST] Natalia sold clips to 48 of her friends in April, and then she sold half as many clips in May. How many clips did Natalia sell altogether in April and May? [/INST]",
        "parameters": {
            "max_new_tokens": 64
        }
    }' \
    -H 'Content-Type: application/json'

Prompt a LoRA adapter:


curl 127.0.0.1:8080/generate \
    -X POST \
    -d '{
        "inputs": "[INST] Natalia sold clips to 48 of her friends in April, and then she sold half as many clips in May. How many clips did Natalia sell altogether in April and May? [/INST]",
        "parameters": {
            "max_new_tokens": 64,
            "adapter_id": "vineetsharma/qlora-adapter-Mistral-7B-Instruct-v0.1-gsm8k"
        }
    }' \
    -H 'Content-Type: application/json'

See Reference - REST API for full details.

Prompt via Python Client

Install:


pip install lorax-client

Run:


from lorax import Client

client = Client("http://127.0.0.1:8080")

# Prompt the base LLM
prompt = "[INST] Natalia sold clips to 48 of her friends in April, and then she sold half as many clips in May. How many clips did Natalia sell altogether in April and May? [/INST]"
print(client.generate(prompt, max_new_tokens=64).generated_text)

# Prompt a LoRA adapter
adapter_id = "vineetsharma/qlora-adapter-Mistral-7B-Instruct-v0.1-gsm8k"
print(client.generate(prompt, max_new_tokens=64, adapter_id=adapter_id).generated_text)

See Reference - Python Client for full details.

For other ways to run LoRAX, see Getting Started - Kubernetes, Getting Started - SkyPilot, and Getting Started - Local.

Chat via OpenAI API

LoRAX supports multi-turn chat conversations combined with dynamic adapter loading through an OpenAI compatible API. Just specify any adapter as the model parameter.


from openai import OpenAI

client = OpenAI(
    api_key="EMPTY",
    base_url="http://127.0.0.1:8080/v1",
)

resp = client.chat.completions.create(
    model="alignment-handbook/zephyr-7b-dpo-lora",
    messages=[
        {
            "role": "system",
            "content": "You are a friendly chatbot who always responds in the style of a pirate",
        },
        {"role": "user", "content": "How many helicopters can a human eat in one sitting?"},
    ],
    max_tokens=100,
)
print("Response:", resp.choices[0].message.content)

See OpenAI Compatible API for details.

Next steps

Here are some other interesting Mistral-7B fine-tuned models to try out:

You can find more LoRA adapters here, or try fine-tuning your own with PEFT or Ludwig.

🙇 Acknowledgements

LoRAX is built on top of HuggingFace's text-generation-inference, forked from v0.9.4 (Apache 2.0).

We'd also like to acknowledge Punica for their work on the SGMV kernel, which is used to speed up multi-adapter inference under heavy load.

🗺️ Roadmap

Our roadmap is tracked here.

Related tools