NVIDIA NeMo AutoModel Accelerates MoE Fine-Tuning 3.7x with Expert Parallelism and DeepEP
NVIDIA has introduced NeMo AutoModel, an open library that builds on Hugging Face Transformers v5, delivering 3.4-3.7x higher training throughput and 29-32% less GPU memory for MoE models via Expert Parallelism, DeepEP fused all-to-all dispatch, and TransformerEngine kernels—using the same from_pretrained() API with only an import change.
NVIDIA
Hugging Face
Moonshot AI
DeepSeek




