New DeepSeek-V3 technical report: how hardware-software co-design enables low-cost training of large models
A new technical paper from the DeepSeek team, with CEO Wenfeng Liang as co-author, explores hardware-aware model co-design to reduce LLM training costs. Using DeepSeek-V3 trained on 2048 NVIDIA H800 GPUs as a case study, the paper details innovations in memory efficiency (MLA), sparse computation (DeepSeekMoE), FP8 training, and interconnect-aware routing.
DeepSeek
NVIDIA
Alibaba/Qwen
Meta




