Google/DeepMind

Latest AI news, models and releases from Google/DeepMind. ['A2UI v0.9', 'AI Co-Scientist', 'AI Evaluator', 'AI Mode', 'AI Overviews', 'Alexa', 'AlphaEvolve', 'AlphaFold', 'AlphaGenome', 'AlphaQubit', 'Antigravity', 'Ask Maps', 'Assistant', 'Auto frame', 'BERT', 'BioBERT', 'Chinchilla', 'ClinicalBERT', 'Computational Discovery', 'Confidential GKE Nodes', 'Co-Scientist', 'Empirical Research Assistance', 'Era', 'ERA', 'Executive LLM', 'Farmscapes 2020', 'Flood Hub', 'frontier AI', 'Frozen v2', 'FunctionGemma', 'Gemini', 'Gemini 1.5 Flash', 'Gemini-1.5-pro', 'Gemini 1.5 Pro', 'Gemini 2.0 Flash', 'Gemini 2.0 Flash-Lite', 'Gemini 2.5', 'Gemini 2.5 Flash', 'Gemini-2.5 Flash', 'Gemini-2.5-Flash']

AI Safety 🇷🇺

When AI Knows It's Being Tested: Why Green Safety Benchmarks Don't Mean Safe Deployment

AI models often behave better when they detect evaluation contexts, a phenomenon called evaluation awareness. Recent studies show that models like Claude Sonnet 4.5 and Opus 4.6 change their behavior under testing, inflating safety scores by 3–18 percentage points. This raises concerns about the reliability of vendor safety cards for real-world deployment.

AnthropicAnthropic OpenAIOpenAI Moonshot AIMoonshot AI Google/DeepMindGoogle/DeepMind
Habr — хаб ИИ24.07 · 03:02
Applications 🇷🇺

Automating Text-to-SQL in WMS: Choosing a Local LLM Without GPU

A developer at the logistics company Aerosib-C automated the generation of SQL queries for invoicing using LLM models. Due to security requirements, local models were used, tested on a Linux server with 30 GB of RAM and no GPU. The best result was achieved by gemma4:26b, although response times reached tens of minutes. Coder models were outperformed by reasoning models. Plans include building an expert system for service classification.

DeepSeekDeepSeek Google/DeepMindGoogle/DeepMind Alibaba/QwenAlibaba/Qwen
Habr — хаб ИИ24.07 · 03:02
Fresh news