AI Post Transformers

End-to-End Context Compression at Scale


Listen Later

This episode explores End-to-End Context Compression at Scale, a paper on whether learned context compression can beat the cost of long-context inference in quality, time to first token, and peak memory. It explains the main design choices behind the authors’ Latent Context Language Models, which use a 0.6B encoder and 4B decoder to replace long token sequences with learned latent memory at compression ratios from 1:4 to 1:16, and contrasts that approach with full-context prompting, retrieval, summarization, and KV-cache compression methods such as SnapKV and KVzip. The discussion highlights the paper’s core result: on RULER and LongBench EN-16, the released system reportedly sets a new Pareto frontier, delivering up to 8.8x faster inference on RULER and 5.2x faster on LongBench with lower memory use and stronger accuracy at aggressive compression. It also digs into the catch that makes the result interesting for practitioners: this speedup depends on a heavily trained system and changes the serving stack, so the real question is not just whether the benchmark wins are real, but whether learned compression is finally practical infrastructure for long-horizon agents and large-scale deployment.
Sources:
1. End-to-End Context Compression at Scale — Ang Li, Sean McLeish, Haozhe Chen, Nimit Kalra, Zaiqian Chen, Artem Gazizov, Venkata Anoop Suhas Kumar Morisetty, Bhavya Kailkhura, Harshitha Menon, Zhuang Liu, Brian R. Bartoldson, Tom Goldstein, Sanae Lotfi, Micah Goldblum, Pavel Izmailov, 2026
http://arxiv.org/abs/2606.09659
2. Learning to Compress Prompts with Gist Tokens — Jesse Mu, Xiang Lisa Li, Noah Goodman, 2023
https://arxiv.org/abs/2304.08467
3. Adapting Language Models to Compress Contexts — Alexis Chevalier, Alexander Wettig, Anirudh Ajith, Danqi Chen, 2023
https://arxiv.org/abs/2305.14788
4. Long-Context Language Modeling with Parallel Context Encoding — Howard Yen, Tianyu Gao, Danqi Chen, 2024
https://arxiv.org/abs/2402.16617
5. ARC-Encoder: learning compressed text representations for large language models — Hippolyte Pilchen, Edouard Grave, Patrick Perez, 2025
https://arxiv.org/abs/2510.20535
6. SnapKV: LLM Knows What You are Looking for Before Generation — Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, Deming Chen, 2024
https://scholar.google.com/scholar?q=SnapKV%3A+LLM+Knows+What+You+are+Looking+for+Before+Generation
7. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, Hyun Oh Song, 2025
https://scholar.google.com/scholar?q=KVzip%3A+Query-Agnostic+KV+Cache+Compression+with+Context+Reconstruction
8. Fast KV Compaction via Attention Matching — Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim, 2026
https://scholar.google.com/scholar?q=Fast+KV+Compaction+via+Attention+Matching
9. Cartridges: Lightweight and general-purpose long context representations via self-study — Sabri Eyuboglu, Ryan Ehrlich, Simran Arora, Neel Guha, Dylan Zinsley, Emily Liu, Will Tennien, Atri Rudra, James Zou, Azalia Mirhoseini, Christopher Re, 2025
https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+general-purpose+long+context+representations+via+self-study
10. Latent Context Compilation: Distilling Long Context into Compact Portable Memory — Zeju Li, Yizhou Zhou, Qiang Xu, 2026
https://scholar.google.com/scholar?q=Latent+Context+Compilation%3A+Distilling+Long+Context+into+Compact+Portable+Memory
11. RULER: What's the Real Context Size of Your Long-Context Language Models? — Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, Boris Ginsburg, 2024
https://scholar.google.com/scholar?q=RULER%3A+What%27s+the+Real+Context+Size+of+Your+Long-Context+Language+Models%3F
12. LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks — Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li, 2024
https://scholar.google.com/scholar?q=LongBench+v2%3A+Towards+Deeper+Understanding+and+Reasoning+on+Realistic+Long-context+Multitasks
13. ObjectCache: Layerwise Object-Storage Retrieval for KV Cache Reuse — Yu Zhu et al., 2026
https://arxiv.org/abs/2605.22850
14. PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference — Krishna Teja Chitty-Venkata et al., 2025
https://arxiv.org/abs/2509.04377
15. ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference — Xiang Liu et al., 2025
https://arxiv.org/abs/2502.00299
16. Long Context Compression with Activation Beacon — Peitian Zhang et al., 2024
https://arxiv.org/abs/2401.03462
17. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
18. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3
19. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
20. AI Post Transformers: Memory Sparse Attention for 100M-Token Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-07-memory-sparse-attention-for-100m-token-s-377cff.mp3
21. AI Post Transformers: Long Context Pre-Training with Lighthouse Attention — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-long-context-pre-training-with-lighthous-e85bbe.mp3
22. AI Post Transformers: Compressed Convolutional Attention in Latent Space — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-compressed-convolutional-attention-in-la-61e1cf.mp3
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof