QuantCode-Bench: a new benchmark for evaluating LLMs' ability to generate executable algorithmic trading strategies
Researchers developed QuantCode-Bench, a benchmark to assess how well large language models (LLMs) can generate executable algorithmic trading strategies from textual specifications. The benchmark evaluates not only code correctness but also semantic alignment with the original trading idea through four stages: compilation, backtest, trade generation, and judge verification.
OpenAI
Anthropic
Google/DeepMind
xAI
DeepSeek
Alibaba/Qwen
Moonshot AI



