llama.cpp

github.com/ggml-org/llama.cpp

Visit Website

A high-performance C++ inference engine for running large language models on CPU and GPU with minimal dependencies, supporting quantized GGUF models for consumer hardware.

Why it is useful

The engine that made local LLMs practical on consumer hardware. Its quantization support lets a 7B model run on 8GB of RAM and a 70B model run on a machine with two mid-range GPUs. Being the most widely supported inference backend means virtually every open-weight model has a GGUF version compatible with it. Other tools like Ollama and LM Studio run llama.cpp under the hood.