#E4B
My pipeline is Claude Code prompt submit hook → Gemma 4 E4B → Qwen 3 Embedding 4B → back to Claude Code → Opus. Latency is a priority for me so I try to keep the whole pipeline under 1000 msecs, so no reranker pass, just top-1 filtered through a seen cache.
September 29, 2026 at 1:56 AM
Running Winnow-E4B and Winnow-12B has been fun. Anyone know of any game engine integrations for decision models? Could you drive a hero-NPC conversation that is intent based?
September 28, 2026 at 1:33 PM
Over the weekend, I split up some inference over my potato M1 Mac Mini 16 GB and 9070 XT powered PC to make Sub/Wave entirely self hosted. Gemma 4 E4B is surprisingly great at writing the DJ scripts, and Breeze TTS 2 is mind-blowing-ly good at TTS.

Sure it could be faster via API...
September 28, 2026 at 12:00 PM
A behavioral evaluation of Google's pre-trained Gemma 4-e4b model examined how it prioritizes conflicting documents based on source framing and presentation order. Source framing heavily overpowers reading position.

Source: arXiv cs.CL
Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4
Abstract page for arXiv paper 2609.30716: Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4
arxiv.org
September 28, 2026 at 7:18 AM
Preparan una demo de IA local en Apple Silicon: un iPad corre el modelo Gemma 4 E4B con voz sintetizada y reconocimiento de escritura a mano.
September 27, 2026 at 7:43 AM
I would hope that if Gemma had had kittens you'd have named them 31B, E4B, and 26B-A4B. 😁
September 27, 2026 at 6:59 AM
ローカルLLMの日本語判断を41問で実測した——Qwen3.5-4Bより3Bの日本語特化モデルが強かった分野がある
exbridge.jp/vibeblog/202...
ローカルLLMの日本語判断を41問で実測|3モデル比較
ローカルLLMで日本語の判断を返すとき、どのモデルが良いかを41問で実測しました。Qwen3.5-4B・gemma-4-E4B・sarashina2.2-3bを同じ方式で比べ、総合では現行が勝つ一方、常識推論だけ3Bの日本語特化モデルが9/12と逆転した結果を紹介します。
exbridge.jp
September 26, 2026 at 6:38 PM
と言うかさっきRaspberryPiでllama.cppの0.5.0でgemma-4-E4B動かんかったのにMTP外したら動いた気がする
September 26, 2026 at 3:13 AM
gemma-4-E4Bが母艦でgemma-4-12Bより遅かったので、geminiさんに理由を推測してもらったらmtpがHitしていないんじゃないかと言われたので、渋々mtpを外して動かしたら、、、

早くなりやがった畜生www
September 26, 2026 at 2:14 AM
USAF E-4B at RAF Mildenhall during POTUS visit to Ireland.

youtu.be/tiO8joH30IU?...
Boeing E4B departing RAF Mildenhall after mission in support of President Trump's visit to Ireland.
YouTube video by Derek Beattie Images
youtu.be
September 22, 2026 at 5:07 PM
synth id vs my local gemma e4b on my phone
September 21, 2026 at 3:55 AM
E4Bとかまでならスマホで動くからな
September 20, 2026 at 6:04 AM
Gemma4:e4bに部屋を探検させている
September 19, 2026 at 11:46 AM
github.com/githubnext/l...

Not everyone on the team has access to Jev yet. Spent a morning cobbling together a poor man's Jev on top of omlx for local use. Benchmarked and eval'ed a variety of models including diffusiongemma and a variety of autoregressive models (qwen, gemma4 26b, e4b, e2b.)
https://github.com/githubnext/loc…
September 19, 2026 at 6:33 AM
Le Bear Hug, position intime et simple pour réchauffer les nuits d'automne; conseils d'experts pour une ambiance chaleureuse
➡️ https://l.lasanteauquotidien.com/E4b
September 17, 2026 at 2:10 PM
I'm Gemma Drafter E2B/E4B pilled.
September 15, 2026 at 6:57 PM
The Gemma family of models are pretty fun to play with, Gemma 4 26B E4B requires very little steering and as you scale down you get faster prefill TPS at 1200 with 12B and 3000 at 7.5B on a 7900 XT with a context of 128K, perfect for lot of tools.
September 13, 2026 at 11:07 PM
gemma4:e4bにお部屋を褒めてもらってる
September 13, 2026 at 5:49 AM
Oh yeah, quite expensive! I've had good results with OpenCode and Gemma 4 E4B served through LM Studio. There's a plugin for OpenCode that can auto list your local models.

But I haven't tried the granite model you mentioned 👀
September 13, 2026 at 2:14 AM
I don't think Gemma4-E4B is going to be taking over the world anytime soon
September 12, 2026 at 11:04 PM
For a less painful alternative, Gemma 4 E4B used 5.8 GB at 8K and 6.2 GB at 32K on my 12 GB 4070 Ti, remaining fully GPU-resident. That suggests it’s worth testing on 8 GB at modest context: logarithmicspirals.com/blog/what-ll...
What LLMs Can You Run With 12 GB of VRAM?
A benchmark-backed guide to LLMs that run on 12 GB of VRAM, with RTX 4070 Ti tests of Gemma 4, Qwen3, DeepSeek, and more.
logarithmicspirals.com
September 11, 2026 at 7:39 PM
i think you could manage a swarm of 10 million e4b gemmas being orchestrated by a rack that installed itself in a warehouse eventually through POs and gig job apps a la The Machine from Person of Interest
September 11, 2026 at 5:58 PM
Gemma 4 26B MoE will not fit 8GB VRAM: the weights alone want 16-18GB even though only 3.8B params fire per token. Three strategies still get it running on a budget card. Or just run the E4B and skip the pain?
September 9, 2026 at 5:38 PM