This episode explores a May 7, 2026 arXiv paper on Lighthouse Attention and asks whether long-context language models can be pretrained cheaply with hierarchical sparse attention, then switched back to standard dense attention late in training without losing dense-model quality. It explains why long-context training is so expensive even with FlashAttention, contrasting dense quadratic attention with sparse and hierarchical schemes that try to narrow which tokens interact. The discussion walks through the paper’s core design: building multi-level pooled query/key/value pyramids, using a gradient-free top-K selector to choose relevant causal subsequences, running ordinary FlashAttention on that smaller set, and scattering the results back. Listeners would find it interesting because it frames the method as a practical systems bet with potentially major implications for 128K- to million-token pretraining, while also stressing that the evidence is still preliminary and far from proving it works at frontier scale.
Sources:
1. Long Context Pre-Training with Lighthouse Attention
https://arxiv.org/pdf/2605.06554
2. H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for Sequences — Zhenhai Zhu, Radu Soricut, 2021
https://scholar.google.com/scholar?q=H-Transformer-1D%3A+Fast+One-Dimensional+Hierarchical+Attention+for+Sequences
3. LongT5: Efficient Text-To-Text Transformer for Long Sequences — Mandy Guo, Joshua Ainslie, David Uthus, Santiago Ontanon, Jianmo Ni, Yun-Hsuan Sung, Yinfei Yang, 2021
https://scholar.google.com/scholar?q=LongT5%3A+Efficient+Text-To-Text+Transformer+for+Long+Sequences
4. HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention — Yufei Xu, Fanxu Meng, Fan Jiang, Yuxuan Wang, Ruijie Zhou, Jiexi Wu, Zhixin Pan, Zhaohui Wang, Xiaojuan Tang, Wenjie Pei, Tongxuan Liu, Di Yin, Xing Sun, Muhan Zhang, 2026
https://scholar.google.com/scholar?q=HISA%3A+Efficient+Hierarchical+Indexing+for+Fine-Grained+Sparse+Attention
5. Long Context Pre-Training with Lighthouse Attention — Bowen Peng, Subho Ghosh, Jeffrey Quesnelle, 2026
https://scholar.google.com/scholar?q=Long+Context+Pre-Training+with+Lighthouse+Attention
6. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention — Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Y. X. Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, Ming Zhang, Wenfeng Liang, Wangding Zeng, 2025
https://scholar.google.com/scholar?q=Native+Sparse+Attention%3A+Hardware-Aligned+and+Natively+Trainable+Sparse+Attention
7. MoBA: Mixture of Block Attention for Long-Context LLMs — Enzhe Lu, Zhejun Jiang, Jingyuan Liu, Yulun Du, Tao Jiang, Chao Hong, Shaowei Liu, Weiran He, Enming Yuan, Yuzhi Wang, Zhiqi Huang, Huan Yuan, Suting Xu, Xinran Xu, Guokun Lai, Yanru Chen, Huabin Zheng, Junjie Yan, Jianlin Su, Yuxin Wu, Neo Y. Zhang, Zhilin Yang, Xinyu Zhou, Mingxing Zhang, Jiezhong Qiu, 2025
https://scholar.google.com/scholar?q=MoBA%3A+Mixture+of+Block+Attention+for+Long-Context+LLMs
8. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Re, 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness
9. Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models — Xiang Hu, Zhanchao Zhou, Ruiqi Liang, Zehuan Li, Wei Wu, Jianguo Li, 2025
https://scholar.google.com/scholar?q=Every+Token+Counts%3A+Generalizing+16M+Ultra-Long+Context+in+Large+Language+Models
10. Hyperattention: Long-context attention in near-linear time — not verified from snippet, recent (not verified from snippet)
https://scholar.google.com/scholar?q=Hyperattention%3A+Long-context+attention+in+near-linear+time
11. When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models — not verified from snippet, recent (not verified from snippet)
https://scholar.google.com/scholar?q=When+Does+Content-Based+Routing+Work%3F+Representation+Requirements+for+Selective+Attention+in+Hybrid+Sequence+Models
12. Delta attention: Fast and accurate sparse attention inference by delta correction — not verified from snippet, recent (not verified from snippet)
https://scholar.google.com/scholar?q=Delta+attention%3A+Fast+and+accurate+sparse+attention+inference+by+delta+correction
13. Spargeattention: Accurate and training-free sparse attention accelerating any model inference — not verified from snippet, recent (not verified from snippet)
https://scholar.google.com/scholar?q=Spargeattention%3A+Accurate+and+training-free+sparse+attention+accelerating+any+model+inference
14. Post-training sparse attention with double sparsity — not verified from snippet, recent (not verified from snippet)
https://scholar.google.com/scholar?q=Post-training+sparse+attention+with+double+sparsity
15. AI Post Transformers: Memory Sparse Attention for 100M-Token Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-07-memory-sparse-attention-for-100m-token-s-377cff.mp3
16. AI Post Transformers: RetrievalAttention for Long-Context LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-17-retrievalattention-for-long-context-llm-ddf566.mp3
17. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
18. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3
19. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
20. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
Interactive Visualization: Long Context Pre-Training with Lighthouse Attention