We made local DiffusionGemma inference 1.8× faster.
Run it on 18GB RAM via Unsloth Studio.
GitHub: github.com/unslothai/un...
Guide: unsloth.ai/docs/models/...
We made local DiffusionGemma inference 1.8× faster.
Run it on 18GB RAM via Unsloth Studio.
GitHub: github.com/unslothai/un...
Guide: unsloth.ai/docs/models/...
unlike previous language diffusion models, this one doesn’t suck, and it’s very fast
blog.google/innovation-a...
unlike previous language diffusion models, this one doesn’t suck, and it’s very fast
blog.google/innovation-a...
Performance is solid: ~35-60 ms for single step latency and batch@32 is ~100-123 requests/sec.
Performance is solid: ~35-60 ms for single step latency and batch@32 is ~100-123 requests/sec.
it is *insane* to see 500 tokens per second at home
it is *insane* to see 500 tokens per second at home
They feel text diffusion models open up a radically different part of the latency–quality Pareto frontier and hope the report makes it easier for researchers and engineers to understand the model, build on it, and create things we haven’t thought of
Here's how I turned DiffusionGemma into a multimodal visual classifier in Java! 🧵👇
Here's how I turned DiffusionGemma into a multimodal visual classifier in Java! 🧵👇
We get Diffusion Gemma!
4x speed up! ⚡⚡⚡💎
blog.google/innovation-a...
We get Diffusion Gemma!
4x speed up! ⚡⚡⚡💎
blog.google/innovation-a...
They feel text diffusion models open up a radically different part of the latency–quality Pareto frontier and hope the report makes it easier for researchers and engineers to understand the model, build on it, and create things we haven’t thought of
They feel text diffusion models open up a radically different part of the latency–quality Pareto frontier and hope the report makes it easier for researchers and engineers to understand the model, build on it, and create things we haven’t thought of
The most cutting-edge model publicly released, DiffusionGemma by Google, doesn't primarily use next-token selection.
The most cutting-edge model publicly released, DiffusionGemma by Google, doesn't primarily use next-token selection.
In theory, denoising tokens in parallel could work better for OCR correction since context is seen upfront?
Pointed it at 19th-century newspaper OCR. It corrected better than the autoregressive baseline — at ~8x the speed.
In theory, denoising tokens in parallel could work better for OCR correction since context is seen upfront?
Pointed it at 19th-century newspaper OCR. It corrected better than the autoregressive baseline — at ~8x the speed.
The new 26B-A4B diffusion text model runs locally on 18GB RAM.
It supports high-speed text generation, thinking, image, video and 256K context.
Run and train via Unsloth Studio.
GGUF: huggingface.co/unsloth/diff...
Guide: unsloth.ai/docs/models/...
The new 26B-A4B diffusion text model runs locally on 18GB RAM.
It supports high-speed text generation, thinking, image, video and 256K context.
Run and train via Unsloth Studio.
GGUF: huggingface.co/unsloth/diff...
Guide: unsloth.ai/docs/models/...