WebLLM

webllm.mlc.ai

Visit Website

In-browser LLM inference powered by WebGPU — runs full language models entirely inside Chrome or Edge with no server, no API key, and no data leaving your device.

Why it is useful

A technically remarkable demo that is also genuinely useful: once the model downloads to your browser cache, it works completely offline with no API costs. Privacy is absolute since computation happens in your GPU through WebGPU. Performance on recent hardware is surprisingly close to server-side inference for 3–7B models. Worth knowing about for privacy-sensitive applications or offline use cases.