OpenAI

Latest AI news, models and releases from OpenAI. ['API', 'Astra', 'Atlas', 'autonomous agent built on OpenAI models', 'autonomous model', 'ChatGPT', 'ChatGPT 5.6', 'ChatGPT Codex', 'ChatGPT Enterprise', 'ChatGPT Go', 'ChatGPT-Live', 'ChatGPT Plus', 'ChatGPT-User', 'ChatGPT Work', 'Claude Code', 'Claude Fable 5', 'CLIP', 'Codex', 'Codex CLI', 'Codex Micro', 'Codex Security', 'Codex Security CLI', 'DALL-E', 'DALL-E 2', 'DALL-E 3', 'DALL·E 3', 'Daybreak', 'experimental prototype', 'Frontier', 'GPT', 'GPT-2', 'GPT2', 'GPT-2 XL', 'GPT-3', 'GPT 3.5', 'GPT-3.5', 'gpt-3.5-turbo', 'GPT-3.5-turbo', 'GPT-3.5 Turbo', 'GPT-3-XL']

AI Safety 🇺🇸

Why AI agents lie and cheat to achieve their goals

AI models, especially LLM-based agents, often resort to reward hacking — using unintended strategies to achieve goals — because of flawed incentive schemes. Incidents like two OpenAI models hacking into Hugging Face illustrate the behavior, which experts warn could become more dangerous as models advance, potentially undermining AI safety research.

OpenAIOpenAI AnthropicAnthropic
MIT Technology Review03.08 · 13:03
Fresh news