On-device inference

Small model, running locally

No API key, no server call at inference time. SmolLM2-135M-Instruct (180MB) runs entirely in this browser tab via WebGPU. Load it once, then try airplane mode (without reloading the page) โ€” extraction still works.

โšก WebGPU accelerated ๐Ÿ”Œ Works with wifi off ๐Ÿ”’ Zero network calls at inference

Loading model...