vLLM
github.com/vllm-project/vllm
A high-throughput production inference engine for large language models using PagedAttention memory management, capable of serving dozens of concurrent users from a single GPU.
Why it is useful
The standard choice for self-hosting LLMs at scale. PagedAttention allows it to handle 20–30 concurrent users on hardware where naive inference would handle one at a time. Provides an OpenAI-compatible API, supports continuous batching for high utilization, and works with virtually every major open-weight model. Used in production by companies that need reliable LLM serving without vendor lock-in.