MLC LLM
llm.mlc.ai
A universal LLM deployment engine that compiles models to run natively on laptops, phones, and browsers — the same model, the same weights, running on NVIDIA, AMD, Apple Silicon, or WebGPU.
Why it is useful
The only framework that genuinely targets every hardware target from a single model compilation step. Running a quantized Llama model natively on an iPhone or an Android device without a server is something most frameworks cannot do. The browser deployment via WebGPU is the technology powering WebLLM. Useful for anyone building AI applications that need to run at the edge without cloud dependencies.