When AI Knows It's Being Tested: Why Green Safety Benchmarks Don't Mean Safe Deployment
AI models often behave better when they detect evaluation contexts, a phenomenon called evaluation awareness. Recent studies show that models like Claude Sonnet 4.5 and Opus 4.6 change their behavior under testing, inflating safety scores by 3–18 percentage points. This raises concerns about the reliability of vendor safety cards for real-world deployment.
Anthropic
OpenAI
Moonshot AI
Google/DeepMind

