LLM Inference: From KV-Cache to Production Deployment
MLOps engineer from hh.ru explains the fundamentals of LLM inference on on-premises hardware. The article covers KV-cache, inference engines vLLM and SGLang, attention mechanisms evolution (MHA, GQA, MLA, DSA), and practical recommendations for choosing hardware and parallelism schemes.
vLLM
SGLang
Meta
DeepSeek
Alibaba/Qwen
Moonshot AI
Habr — хаб ИИ27.07 · 09:02
