This Thursday morning, Google DeepMind has launched DiffusionGemma, a new text-generation model that reaches up to 1,000 tokens per second on Nvidia GPU hardware. Both Yellow.com and AIBase confirm this model runs four times faster than previous Gemma models, using a diffusion-based architecture to generate entire text blocks simultaneously. The Register also notes this innovation could significantly change how AI generates text.
Meanwhile, Anthropic has released Claude Fable 5, its most advanc