Run real AI models directly in your browser — now including text generation. Model weights download once from Hugging Face and inference happens on your device: free, unlimited, private. No API key, no server, no rate limits.
PER person ·
ORG organization ·
LOC location ·
MISC other
Why run models in the browser? Hosted inference APIs meter usage and the free tiers keep shrinking — GitHub's own Models inference API was retired in 2026. Local inference is the durable alternative: the model runs on your own hardware, so it costs nothing, cannot be rate-limited, and your prompts never leave the page. A 135M instruct model is genuinely useful for short drafts and lists; a 0.5B model trades a heavier download for a bit more coherence. Both are real transformer models with chat templates, not toys.