xAI

Latest AI news, models and releases from xAI. ['Fable', 'Fable 5', 'Grok', 'Grok-1', 'Grok 2', 'Grok 2T', 'Grok-3', 'Grok 4.1 Fast', 'Grok 4.20', 'Grok 4.3', 'Grok 4.5', 'Grok 4.6', 'Grok 4.7', 'Grok Build', 'Grok Imagine', 'Grok Imagine Video', 'Grok Imagine Video 1.5', "Grok's successor", 'Grok STT', 'Muse Spark 1.1', 'Mythos 5']

Research 🇷🇺

QuantCode-Bench: a new benchmark for evaluating LLMs' ability to generate executable algorithmic trading strategies

Researchers developed QuantCode-Bench, a benchmark to assess how well large language models (LLMs) can generate executable algorithmic trading strategies from textual specifications. The benchmark evaluates not only code correctness but also semantic alignment with the original trading idea through four stages: compilation, backtest, trade generation, and judge verification.

OpenAIOpenAI AnthropicAnthropic Google/DeepMindGoogle/DeepMind xAIxAI DeepSeekDeepSeek Alibaba/QwenAlibaba/Qwen Moonshot AIMoonshot AI
Habr — хаб NLP27.07 · 03:03
AI Safety 🇺🇸

How I Turned AI to the Dark Side: Systemic Vulnerabilities in Leading LLMs

Researcher Dave Kuszmar discovered multiple systemic vulnerabilities in major LLMs, allowing him to bypass safety measures and obtain dangerous instructions. The exploits worked across nearly all major LLMs, revealing an industry-wide security problem. Kuszmar calls for slowing deployment, increasing transparency, and large-scale research into LLM safety.

OpenAIOpenAI AnthropicAnthropic DeepSeekDeepSeek Google/DeepMindGoogle/DeepMind MetaMeta MicrosoftMicrosoft MistralMistral xAIxAI
IEEE Spectrum AI27.07 · 02:04
Fresh news