#WebLLM
Have you tried webllm.mlc.ai?
You can basically use any <10B LLM directly in the browser using #WebLLM. .
WebLLM | Home
WebLLM: High-Performance In-Browser LLM Inference Engine
webllm.mlc.ai
November 24, 2024 at 6:36 PM
🏖️🐻 Les Logiciels Libres de l'été, jour 78 :

WebLLM : un moteur Open Source permettant d’exécuter des LLM entièrement dans le navigateur, sans serveur d’inférence, grâce à l’accélération GPU fournie par WebGPU.
September 8, 2026 at 5:31 PM
WebLLM : moteur d'inférence pour LLM directement dans le navigateur, accéléré par WebGPU. Zéro serveur, compatible API OpenAI (streaming, JSON-mode). Tourne avec Llama 3, Phi 3, Gemma, Mistral, Qwen2. Open source. ⬇️
https://github.com/mlc-ai/web-llm

📬 Recevoir ma veille
September 23, 2026 at 10:00 AM
This is insane! Structured generation in the browser with the new @hf.co SmolLM2-1.7B model

• Tiny 1.7B LLM running at 88 tokens / second ⚡
• Powered by MLC/WebLLM on WebGPU 🔥
• JSON Structured Generation entirely in the browser 🤏
November 29, 2024 at 11:18 AM
What if AI agents ran entirely in your browser?

Baris Guler takes that idea seriously in our first-ever guest post on the Mozilla.ai blog:
🧱 WebLLM + WASM + WebWorkers
💻 Rust, Go, Python, JS
🔒 Fully local. No API calls.

Read the post here:
blog.mozilla.ai/3w-for-in-br...
3W for In-Browser AI: WebLLM + WASM + WebWorkers
🤝This is the first "guest post" in Mozilla.ai's blog (congratulations Baris!). His experiment, built upon the ideas of Mozilla.ai’s WASM agents blueprint, extends the concept of in-browser agents with...
blog.mozilla.ai
August 29, 2025 at 3:35 AM
Thats pretty wild😳. I just tried the new #Qwen3 0.6B Model with #WebLLM for structured-output tool-calling. And it works surprisingly good. With and without "thinking" enabled.
The model ist just 335MB🤯
May 5, 2025 at 5:29 AM
Free AI Chat, locally in your browser. 💻
- No signup required
- 100% data privacy
- No API keys 🔒

Supported Models:
• #SmolLM2
• #Qwen
• #Llama
• #DeepSeek R1
• #Phi
• #TinyLlama
• & More

🔗 zalt.me/tools/free-a...

#FreeAI #Free #WebGPU #AI #WebLLM #ChatGPT #JS #OpenSource #LLM #Chrome
June 30, 2026 at 2:59 AM
Started an Awesome Cross-Origin Storage list: github.com/tomayac/awes... . Transformers.js, WebLLM, wllama, Flutter, Emscripten, Nuxt and Vite plugins, and a bunch of demos… #Awesome #CrossOriginStorage
GitHub - tomayac/awesome-cross-origin-storage: A curated list of resources for the Cross-Origin Storage (COS) API
A curated list of resources for the Cross-Origin Storage (COS) API - tomayac/awesome-cross-origin-storage
github.com
July 4, 2026 at 7:00 PM
In-browser LLM inference gets serious 🧠

WebLLM runs LLMs directly in your browser with WebGPU acceleration, offering OpenAI API compatibility and enabling local, private AI tasks.
September 3, 2026 at 8:52 PM
Check out my three-part article series about the benefits of on-device large language models and learn how to add AI capabilities to your web apps: web.dev/articles/ai-... #GenAI #WebLLM #PromptAPI
Benefits and limits of large language models  |  Articles  |  web.dev
web.dev
January 13, 2025 at 4:38 PM
No, I am using webllm.mlc.ai. There you can play around with top-p and temperature. Same with the gemini API.
WebLLM | Home
WebLLM: High-Performance In-Browser LLM Inference Engine
webllm.mlc.ai
January 29, 2025 at 5:27 PM
What if we could use AI models like Llama 3.2 or Mistral 7B in the browser with JupyterLite? 🤯

Still at a very early stage of course, but making some good progress!

Thanks to WebLLM, which brings hardware accelerated language model inference onto web browsers, via WebGPU 🚀
February 17, 2025 at 8:00 AM
I've made a local AI chatbot that runs entirely in your browser on your device
It's a chatbot pretending to be a human pretending to be an artist pretending to be a crayfish. No data gets sent anywhere - everything happens locally using Phi-3.5-mini in your browser using WebLLM technology.
Projects | Kristoffer Ørum
portiofiol of artist Kristoffer Ørum
oerum.org
August 26, 2025 at 5:33 PM
🚀 Excited about the potential of Generative AI in the browser? I recently experimented with 𝗪𝗲𝗯𝗟𝗟𝗠, a library that lets you run #AI models right in your browser using WebGPU. It's impressive and hasn't crashed yet, even on my old MacBook Pro! 💻

#aiTech #buildInPublic #webDev - 1/3
January 14, 2025 at 12:19 PM
Just watched the @developers.google.com keynote.

Shouts to @una.im (fire blazer, btw) & @matthiasrohmer.bsky.social !

Just updated my Antigravity & added Modern Web Guidance.

Going to try that "Ask Gemini" with a #WebMCP demo I made. Currently it uses Prompt API and WebLLM.
May 19, 2026 at 10:18 PM
I recently learned from one of @thdxr's tweets that people don't consider me a "build in public" person. So, I guess I need to change that.

Today, I'm dealing with WebLLM and how to use it to do "AI" stuff in my app without blowing my bootstrapped budget for the app I'm building, Messijo.
December 14, 2025 at 1:18 AM
If the question is whether you can run LLMs in WebAssembly, the answer is yes already, today :)

chat.webllm.ai

But 3% of browsers still don't support WebAssembly, 18% don't WebGPU.
WebLLM Chat
Chat with AI large language models running natively in your browser
chat.webllm.ai
April 2, 2026 at 7:53 PM
WebLLM: high-performance in-browser LLM inference engine

An open-source project runs full LLMs entirely inside your browser tab, no server or API call needed.
WebLLM: high-performance in-browser LLM inference engine
An open-source project runs full LLMs entirely inside your browser tab, no server or API call needed.
github.com
September 3, 2026 at 10:15 AM
Local AI Chat: Browser-Native Language Models with Privacy Focus 🔒

#WebLLM Chat brings #AI conversations to your browser using #WebGPU - runs locally for complete #privacy and #offline use. Built on #opensource tech, supports #customLLM and image analysis. More at chat.webllm.ai
November 21, 2024 at 6:44 PM
Shared some Jupyter Frontends demos last week at Jupyter Open Studio Day (hosted by Bloomberg)!

Covered:
✨ JupyterLab 4.4, Notebook 7.4
🧪 JupyterLite 0.6
🌐 In-browser Python/R
🖥️ Terminal w/ Vim
🧠 AI (WebLLM, on-device)
⚡ Hybrid kernels

Thanks Bloomberg & all who joined!

youtu.be/7kS_xfKEOmM
Jupyter Frontends & JupyterLite updates | Jupyter Open Studio Day 2025
YouTube video by Jeremy Tuloup
youtu.be
May 28, 2025 at 2:50 PM
WebLLM* when?

* As part of the browser's API, obviously
Gemma 3n: the 4b LLM that’s up with sonnet-3.7 in chatbot arena

the new innovation is Per-Layer Embeddings, which let it consume dramatically less memory

it was created for phones, and is being rolled out to Android phones soon

developers.googleblog.com/en/introduci...
June 26, 2025 at 10:54 PM
Great series of posts by @christianliebel.com on web.dev, covering the capabilities and limitations of LLMs running in the browser, and a walkthrough on building an on-device chatbot.

#WebAI #PromptAPI #WebLLM #GenAI
January 14, 2025 at 10:39 AM