Google Built a Faster AI by Making It Worse. When Is That a Good Trade?
DiffusionGemma, dFlash and TurboQuant get sold as one 6x speedup. They're three different tradeoffs, and collapsing them hides the one that matters.
Originally published elsewhere.

DiffusionGemma is Google’s new diffusion text model. The speed is real. The way it gets sold is not.
On 10 June 2026 Google released DiffusionGemma. Diffusion text models existed before it, from startups and in closed previews. This is the first one you can download from a major lab. It ships under Apache 2.0, and it is fast. Google reports over 1,000 tokens per second on a single H100 and over 700 on a consumer RTX 5090, up to four times the throughput of a comparable autoregressive model.
That is the headline. It is true. It is also where a tempting story starts.
The story goes like this. Pair DiffusionGemma with dFlash and TurboQuant and RAG, wire it into a local document chatbot, and you get one pipeline that hands you a free six times improvement. It is the kind of story that travels well. Three of those four names are real and genuinely good engineering. The problem is that they do three different jobs with three different tradeoff profiles, and folding them into a single number throws away the only information that matters when you decide whether to ship any of them.
So here is the version with the tradeoffs left in.
Three different speedups
DiffusionGemma changes how text is generated. An autoregressive model writes one token at a time, each conditioned on the last. DiffusionGemma starts from a block of 256 placeholder tokens and refines them all in parallel over a handful of denoising passes. On a single-user local box this is a good trade, because your GPU spends most of its time idle between keystrokes anyway. The cost is quality. Google is unusually direct about this. By its own account DiffusionGemma trails standard Gemma 4 on every benchmark Google published for it. This is a lossy speedup. You are paying in output quality for throughput.

dFlash does the opposite trade. It is a speculative decoding method. A small draft model proposes a block of tokens, your normal autoregressive target model verifies them in one pass, and only accepted tokens are kept. The draft happens to be a block-diffusion model, which is where the confusion creeps in, but the generator you actually trust is still your autoregressive target. The output is identical to running the target alone. This is a lossless speedup. You pay nothing in quality. You pay in the engineering to run a second model, and in throughput that tapers as context grows.
TurboQuant changes how vectors are stored. It is a vector quantisation method from Google and NYU that compresses high-dimensional vectors with near-optimal distortion and no training. Its headline use is shrinking the key-value cache by roughly six times. In a RAG pipeline it is just as useful for compressing the embedding index, so retrieval is smaller and faster. That is a memory and search speedup. It does not make your model read more tokens.
Notice what just happened. Lossy generation, lossless decoding, and vector compression are three orthogonal axes. You can adopt any one of them without the others. The popular framing collapses them into a single “six times” figure, which is the KV-cache number from TurboQuant borrowed to describe a context gain it has nothing to do with. Eight thousand characters becoming thirty-two thousand is four times, not six. Characters are not tokens. A smaller RAG index does not extend a context window. Those are three different claims, not one number.

The demo does not run the model in the headline
Here is the detail that should make you slow down. The demo is titled around DiffusionGemma, but its generation path loads a draft model and calls a speculative decoding loop. That is dFlash. The denoising generation that makes DiffusionGemma interesting is not in the hot path at all. The pipeline is named after the model it does not use.
This is not pedantry. If you copied that code expecting diffusion-speed generation and instead got autoregressive speculative decoding, your latency profile, your quality profile and your hardware sizing are all different from what the title promised. Reading the code beats reading the title.

While we are here. The same demo is sold as better OCR, then extracts text with PyPDF. PyPDF reads the embedded text layer of a PDF. It is not OCR. Hand it a scanned image and it returns nothing. DiffusionGemma does accept image input, but the pipeline shown never uses it. So the label oversells the implementation on two separate counts.
“Diffusion cannot self-correct” is the wrong sentence
The most interesting claim attached to these models is that diffusion cannot self-correct, offered as the structural weakness behind wobbly multi-turn behaviour. The instinct is right. The sentence is wrong, and the gap between the two is the whole point.
Within a block, DiffusionGemma self-corrects by design. Bidirectional attention lets every token attend to every other, and the sampler can re-noise low-confidence positions and try again. That is error correction, and it is the mechanism Google leans on for constrained problems like Sudoku.
The real limit is one block out. To generate beyond 256 tokens, DiffusionGemma commits a finished block to the KV cache and starts a fresh canvas conditioned on it. Once committed, a block is frozen. The model can refine what it is drafting now. It cannot revise what it already locked. So the accurate statement is not “cannot self-correct.” It is “cannot revise a committed block.” That predicts a specific failure. In a long answer that spans several blocks, the model can polish the block in front of it but cannot go back and fix an earlier one once it is locked. Whether that compounds across a full multi-turn conversation is a further question, and the block mechanism alone does not answer it. What the mechanism does give you is a place to look.
Why this is an eval problem, not a leaderboard problem
A leaderboard will show you DiffusionGemma sitting below Gemma 4 and let you conclude it is worse. That is the wrong conclusion and the wrong tool. The benchmark gap is real signal, not noise. It tells you a quality cost exists. What it cannot tell you is whether that cost lands anywhere near the work you actually do. So the question is not which model wins a benchmark. The question is which failure you can tolerate for which workload.
For in-line editing, code infilling and rapid local iteration, where you will read and correct the output anyway, lossy parallel generation at four times the speed is a fine trade. For a single-user assistant on a quiet GPU, the throughput is close to free. That qualifier carries weight. The speed edge holds at batch sizes of roughly one to eight, which is exactly single-user, self-hosted territory. Push to batch thirty-two and beyond and ordinary autoregressive models reclaim the lead through KV-cache reuse, so the win does not follow you onto a busy shared server.

This is a local-inference story, and it is worth being honest that it stays one. For quality-critical production text, Google’s own advice is to stay on autoregressive Gemma 4, and you should take it. For a RAG system you actually care about, the part worth keeping from that demo is the lossless one. Speculative decoding for speed without a quality hit, and quantised vectors for a smaller index.
None of that is visible on a leaderboard. All of it is visible the moment you write down what each piece costs. Speed is the easy number to publish. The tradeoff is the number worth knowing.