From DeepSeek V3 to V3.2: Architecture, Sparse Attention, and RL Updates
DeepSeek released V3.2, an open-weight flagship model with performance comparable to GPT-5 and Gemini 3.0 Pro. The model introduces DeepSeek Sparse Attention (DSA) for efficiency, and evolves from a dedicated reasoning model (R1) to a hybrid reasoning model (V3.1, V3.2). The timeline includes V3, R1, R1-0528, V3.1, V3.2-Exp, and V3.2, with architectural highlights like Multi-Head Latent Attention (MLA) and RLVR training.
DeepSeek





