A Visual Guide to Attention Variants in Modern LLMs
This article provides a visual overview of attention mechanisms in modern large language models, covering from standard multi-head attention (MHA) to grouped-query attention (GQA), multi-head latent attention (MLA), sparse attention, and hybrid architectures. It is accompanied by an LLM architecture gallery with 45 entries and poster versions.
OpenAI
Allen Institute for AI
Meta
Alibaba/Qwen
Google/DeepMind
Mistral
Hugging Face
Cohere


