Alibaba Cloud: Zhenwu M890 Chip Adapts to Qwen3.8, Model Goes Live on Bailian for Inference
Alibaba/Qwen
Alibaba Cloud announced that its Zhenwu M890 supernode chip has successfully adapted the Qwen3.8 model, enabling inference on the Bailian platform. The Zhenwu M890 is a new-generation AI chip from T-Head, supporting FP32 to FP4 precision and achieving 800GB/s interconnect via ICN Switch 1.0, allowing a single instance to run a 2.4-trillion-parameter model.
Alibaba Cloud announced on July 23 that the Zhenwu M890 supernode chip has successfully adapted Qwen3.8, Alibaba's flagship model with 2.4 trillion parameters, and is now providing inference services on the Bailian platform. This marks the first time a Chinese supernode has run a model exceeding 2 trillion parameters. The Zhenwu M890, developed by T-Head, is a new-generation AI chip supporting data precisions from FP32 to FP4, suitable for high-precision training and low-precision inference. Through ICN Switch 1.0, 64 Zhenwu M890 chips achieve 800GB/s interconnect, providing 9TB of video memory and enabling expert parallelism for large MoE models. Alibaba also optimized the full stack from chip to cloud for Qwen3.8, improving inference efficiency by up to 1.5x in agentic scenarios.
- Сокращения
- ICN = Interconnect Network — интерконнектная сеть
- GPU = Graphics Processing Unit — графический процессор
- FP32 = Floating Point 32-bit — 32-битная плавающая точка
- FP4 = Floating Point 4-bit — 4-битная плавающая точка
- MoE = Mixture of Experts — смесь экспертов
Source: QbitAI 量子位 —
original
