Google/DeepMind

Latest AI news, models and releases from Google/DeepMind. ['A2UI v0.9', 'AI Co-Scientist', 'AI Evaluator', 'AI Mode', 'AI Overviews', 'Alexa', 'AlphaEvolve', 'AlphaFold', 'AlphaGenome', 'AlphaQubit', 'Antigravity', 'Auto frame', 'BERT', 'BioBERT', 'Chinchilla', 'ClinicalBERT', 'Computational Discovery', 'Confidential GKE Nodes', 'Co-Scientist', 'Empirical Research Assistance', 'Era', 'ERA', 'Executive LLM', 'Farmscapes 2020', 'Flood Hub', 'frontier AI', 'Frozen v2', 'FunctionGemma', 'Gemini', 'Gemini 1.5 Flash', 'Gemini-1.5-pro', 'Gemini 1.5 Pro', 'Gemini 2.0 Flash', 'Gemini 2.0 Flash-Lite', 'Gemini 2.5', 'Gemini 2.5 Flash', 'Gemini-2.5 Flash', 'Gemini-2.5-Flash', 'gemini-2.5-flash-image', 'gemini-2.5-flash-lite']

Mini-PC on Strix Halo under parallel load: 236 tok/s on 32 concurrent requests and three errors

A Beelink GTR9 Pro with Ryzen AI Max+ 395 achieved 236 tok/s aggregate on 32 concurrent requests in short runs, sustaining 226 tok/s average over 30 minutes without throttling. The author discovered a reproducible throughput drop between 8 and 10 concurrent clients across all tested models, and found that speculative decoding (MTP) accelerated single-client performance by 42% but became a penalty under concurrency.

Google/DeepMindGoogle/DeepMind
Habr — хаб NLP24.07 · 10:04
Fresh news