The briefing
Google introduced experimental Multi‑Token Prediction drafters for its Gemma 4 open models, using speculative decoding to anticipate future tokens. Ars Technica reports the approach can significantly speed up generation on edge hardware compared with standard decoding.
Original headline
Google's Gemma 4 AI models get 3x speed boost by predicting future tokens