Anthropic

Latest AI news, models and releases from Anthropic. ['anthropic-ai', 'Anthropic AI Models', 'Claude', 'Claude 2', 'Claude 2.1', 'Claude 3', 'Claude 3.5 Sonnet', 'Claude-3.5-Sonnet', 'Claude 3.5 Sonnet v2', 'Claude 3.7 Sonnet', 'Claude 3 Haiku', 'Claude 3 Opus', 'Claude3 Opus', 'Claude 3 Sonnet', 'Claude 4', 'Claude 4.6', 'Claude 4 Opus', 'Claude 4 Sonnet', 'Claude 5', 'Claude Agent SDK', 'Claude.ai', 'Claude AI', 'ClaudeBot', 'Claude Code', 'Claude Code (Opus 4.7)', 'Claude Code Sonnet 5', 'Claude Cowork', 'Claude Design', 'Claude Desktop', 'Claude Fable', 'Claude Fable 5', 'Claude Fable5', 'Claude Fable 5 Max', 'Claude Haiku', 'Claude Haiku 4.5', 'Claude Instant', 'Claude (language model)', 'Claude Max', 'Claude Mythos', 'Claude Mythos 5']

Agents 🇺🇸

ScarfBench: A Benchmark for Evaluating AI Agents in Enterprise Java Framework Migration

ScarfBench is an open benchmark for evaluating AI agents on cross-framework migration tasks in Enterprise Java, covering Spring, Jakarta EE, and Quarkus. It tests whether migrated applications build, deploy, and preserve behavior. Current agents show overconfidence and struggle with environment issues, with configuration dominating migration effort.

IBMIBM AnthropicAnthropic
Hugging Face blog27.07 · 09:05
Fresh news