how this site works

This website uses WebGPU to run a small language model in your browser. Curated data and rails are built in to help the model stay on track and give helpful answers. There's no backend, so nothing you type leaves the page.

models

Llama-3.2-1B is a small model, taking around a GB of memory, but fast enough for most computers to handle. Gemma-2-2B is double the size and near the limit of what can comfortably run in a browser. Phones get SmolLM2-360M, a lighter model that fits inside mobile Safari's memory limits.

hosting

This site is a simple page stored in S3 and served by CloudFront, costing around $0.50 a month to host. The model weights are stored in S3 too, which keeps them versioned alongside the site.

local · vram — · tps — · pp —