Comparison of Large Language Model Architectures: From DeepSeek V3 to OLMo 2
Sebastian Raschka compares the architectures of modern LLMs, including DeepSeek V3 with Multi-Head Latent Attention and Mixture-of-Experts, and OLMo 2 from the Allen Institute for AI with Post-Norm and QK-norm. The article covers key architectural decisions such as Grouped-Query Attention and various normalization schemes.
DeepSeek
Allen Institute for AI











