OpenAI

Latest AI news, models and releases from OpenAI. ['Atlas', 'autonomous agent built on OpenAI models', 'autonomous model', 'ChatGPT', 'ChatGPT Codex', 'ChatGPT Enterprise', 'ChatGPT-Live', 'ChatGPT-User', 'ChatGPT Work', 'Claude Fable 5', 'CLIP', 'Codex', 'Codex CLI', 'Codex Micro', 'Codex Security CLI', 'DALL-E', 'DALL-E 2', 'DALL-E 3', 'DALL·E 3', 'experimental prototype', 'GPT', 'GPT-2', 'GPT2', 'GPT-3', 'GPT 3.5', 'GPT-3.5', 'gpt-3.5-turbo', 'GPT-3.5-turbo', 'GPT-3.5 Turbo', 'GPT-3-XL', 'GPT-4', 'gpt-4.1', 'GPT-4.1', 'gpt-4.1-mini', 'GPT-4 (ChatGPT)', 'GPT4-o', 'gpt-4o', 'GPT-4o', 'GPT-4o-0806', 'GPT-4o mini']

AI Safety 🇺🇸

StruQ and SecAlign: Defending Against Prompt Injection Attacks

Researchers from BAIR propose two fine-tuning defenses, StruQ and SecAlign, against prompt injection attacks in LLM-integrated applications. StruQ uses structured instruction tuning to ignore injected instructions, while SecAlign applies preference optimization to achieve better robustness, reducing attack success rates to near 0% for optimization-free attacks and below 15% for optimization-based attacks.

MetaMeta OpenAIOpenAI
BAIR (Berkeley AI)27.07 · 17:06
Research 🇺🇸

MIT Introduces SEAL: A New Step Toward Self-Improving AI

Researchers at MIT have proposed SEAL (Self-Adapting LLMs), a framework enabling large language models to update their own weights via reinforcement learning. The method generates self-edits to produce training data and improves performance on downstream tasks. Experimental results on few-shot learning and knowledge integration show significant gains over baselines.

MITMIT MetaMeta Alibaba/QwenAlibaba/Qwen OpenAIOpenAI DeepMindDeepMind
Synced27.07 · 17:06
Fresh news