跳到正文
Ahead of AI· Sebastian Raschka, PhD·· 2026-05-16AI 评分40

Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention

Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention

AI 导读

Google 在 Gemma 4 中引入了跨层 KV 缓存共享机制,通过在不同层之间复用键值投影来减少长上下文推理的内存和计算需求。Gemma 4 E2B 使用 MQA 和滑动窗口注意力,以 4:1 比例优化注意力计算。该架构设计旨在提升推理效率,适用于移动端和嵌入式设备。

来源:Ahead of AI · magazine.sebastianraschka.com