从 SWA 的固定窗口出发,解释 Lightning Indexer 如何为完整前缀打分并选出 top-k、融合式 Sparse Attention 算子如何按索引读取 KV,以及 DSA 如何降低 Attention 的计算量与显存读取量。
Blog
Engineering, reflections, and reviews.
Writing activity
4 posts across 2 active months · last 12 months
- August 2025: 0 posts
- September 2025: 0 posts
- October 2025: 0 posts
- November 2025: 0 posts
- December 2025: 0 posts
- January 2026: 0 posts
- February 2026: 0 posts
- March 2026: 0 posts
- April 2026: 0 posts
- May 2026: 0 posts
- June 2026: 1 post
- July 2026: 3 posts
Less More
从 MHA、MQA、GQA 的 KV head 共享关系出发,解释 MLA 如何通过低秩 KV latent、矩阵吸收和解耦 RoPE 压缩推理缓存。
用 column-wise 和 row-wise 两种切分方式解释 Transformer TP:中间分片如何继续本地计算,输出贡献又如何通过 all-reduce 合并。
从 EngineCore step 的异步执行模型出发,梳理 vLLM 在 PD 分离场景下与 NIXL KV Cache 传输协同的关键路径。