Google/DeepMind

Latest AI news, models and releases from Google/DeepMind. ['A2UI v0.9', 'AI Co-Scientist', 'AI Evaluator', 'AI Mode', 'AI Overviews', 'Alexa', 'AlphaEvolve', 'AlphaFold', 'AlphaGenome', 'AlphaQubit', 'Antigravity', 'Auto frame', 'BERT', 'BioBERT', 'Chinchilla', 'ClinicalBERT', 'Computational Discovery', 'Confidential GKE Nodes', 'Co-Scientist', 'Empirical Research Assistance', 'Era', 'ERA', 'Executive LLM', 'Farmscapes 2020', 'Flood Hub', 'frontier AI', 'Frozen v2', 'FunctionGemma', 'Gemini', 'Gemini 1.5 Flash', 'Gemini-1.5-pro', 'Gemini 1.5 Pro', 'Gemini 2.0 Flash', 'Gemini 2.5', 'Gemini 2.5 Flash', 'Gemini-2.5 Flash', 'Gemini-2.5-Flash', 'gemini-2.5-flash-image', 'Gemini 2.5 Pro', 'Gemini-2.5 Pro']

AI Safety 🇺🇸

AI models prone to 'overthinking' creates security vulnerability

Researchers from Zhejiang University and Alibaba have demonstrated a method to deliberately induce 'overthinking' in reasoning AI models by subjecting them to logically inconsistent prompts. This evolutionary prompt attack can cause outputs up to 26 times longer, effectively acting as a denial-of-service attack on commercial AI services.

DeepSeekDeepSeek Alibaba/QwenAlibaba/Qwen OpenAIOpenAI Google/DeepMindGoogle/DeepMind
IEEE Spectrum AI27.07 · 04:05
Research 🇷🇺

Finam AI Lab Updates Financial Benchmark FINESSE-Bench for LLMs

Finam's Artificial Intelligence Laboratory has released an updated version of the financial benchmark FINESSE-Bench. The new version adds a technical analysis dataset CFTe-like Level 1, fixes 169 CFA-like Level 1 questions, improves metric calculation with bootstrapping, and expands the model pool to 33 in comparative tables.

AnthropicAnthropic Moonshot AIMoonshot AI Google/DeepMindGoogle/DeepMind MiniMaxMiniMax OpenAIOpenAI MetaMeta DeepSeekDeepSeek MistralMistral
Habr — хаб NLP27.07 · 04:05
Fresh news