AI Post Transformers

Vistara Brings CXL Memory to Hyperscale


Listen Later

This episode explores whether CXL memory expansion has finally become practical for hyperscale production, using the 2025 Vistara system as a case study. It explains the core ideas behind CXL, tiered memory, and memory disaggregation, then argues that the real comparison is not against swap but against transparent page placement that keeps hot data in local DRAM and colder pages in a slower expanded tier. The discussion highlights Vistara’s full-stack design, from a custom low-latency ASIC and Linux support to workload-specific tuning on a production server with 768 GB of local DDR5 and 256 GB of CXL-attached DDR4. Listeners would find it interesting because the episode moves past industry hype and examines the concrete tradeoffs around latency, bandwidth, operational complexity, and whether memory can finally be managed as a flexible datacenter resource rather than a fixed property of a single machine.
Sources:
1. Vistara Brings CXL Memory to Hyperscale
https://aisystemcodesign.github.io/papers/isca26/vistara_camera_ready.pdf
2. Software-Defined Far Memory in Warehouse-Scale Computers — H. Andres Lagar-Cavilla, Junwhan Ahn, Suleiman Souhlal, Neha Agarwal, Junaid Shahid, Greg Thelen, Parthasarathy Ranganathan, and others, 2019
https://scholar.google.com/scholar?q=Software-Defined+Far+Memory+in+Warehouse-Scale+Computers
3. TMO: Transparent Memory Offloading in Datacenters — Johannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang, Hao Wang, Blaise Sanouillet, Bikash Sharma, Tejun Heo, Mayank Jain, Chunqiang Tang, Dimitrios Skarlatos, 2022
https://scholar.google.com/scholar?q=TMO%3A+Transparent+Memory+Offloading+in+Datacenters
4. Pond: CXL-Based Memory Pooling Systems for Cloud Platforms — Huaicheng Li, Daniel S. Berger, Stanko Novakovic, Lisa Hsu, Dan Ernst, Pantea Zardoshti, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D. Hill, Marcus Fontoura, Ricardo Bianchini, 2023
https://scholar.google.com/scholar?q=Pond%3A+CXL-Based+Memory+Pooling+Systems+for+Cloud+Platforms
5. TPP: Transparent Page Placement for CXL-Enabled Tiered-Memory — Hasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner, Niket Agarwal, Pallab Bhattacharya, Chris Petersen, Mosharaf Chowdhury, Shobhit Kanaujia, Prakash Chauhan, 2023
https://scholar.google.com/scholar?q=TPP%3A+Transparent+Page+Placement+for+CXL-Enabled+Tiered-Memory
6. Managing Memory Tiers with CXL in Virtualized Environments — Yuhong Zhong, Daniel S. Berger, Carl Waldspurger, Richard Wee, Ishan Agarwal, Raghav Agarwal, Fred Hady, K. Kumar, Mark D. Hill, Mosharaf Chowdhury, Ahmed Cidon, 2024
https://scholar.google.com/scholar?q=Managing+Memory+Tiers+with+CXL+in+Virtualized+Environments
7. Demystifying CXL Memory with Genuine CXL-Ready Systems and Devices — Yongjun Sun, Ye Yuan, Ziming Yu, Ryan Kuper, Chao Song, Jiyong Huang, Honggyu Ji, Saurabh Agarwal, Jingwen Lou, Inhwan Jeong, Rui Wang, Joonho H. Ahn, Tianyin Xu, Nam Sung Kim, 2023
https://scholar.google.com/scholar?q=Demystifying+CXL+Memory+with+Genuine+CXL-Ready+Systems+and+Devices
8. M5: Mastering Page Migration and Memory Management for CXL-based Tiered Memory Systems — Yongjun Sun, Jiyoon Kim, Ziming Yu, Joonwon Zhang, Sangho Chai, Minjae J. Kim, Hyeonsu Nam, Jihye Park, Euna Na, Ye Yuan, Rui Wang, Joonho H. Ahn, Tianyin Xu, Nam Sung Kim, 2025
https://scholar.google.com/scholar?q=M5%3A+Mastering+Page+Migration+and+Memory+Management+for+CXL-based+Tiered+Memory+Systems
9. Dissecting CXL Memory Performance at Scale: Analysis, Modeling, and Optimization — approx. recent systems authors, 2024/2025
https://scholar.google.com/scholar?q=Dissecting+CXL+Memory+Performance+at+Scale%3A+Analysis%2C+Modeling%2C+and+Optimization
10. Improving Key-Value Cache Performance with Heterogeneous Memory Tiering: A Case Study of CXL-Based Memory Expansion — approx. recent cache/tiering authors, 2024/2025
https://scholar.google.com/scholar?q=Improving+Key-Value+Cache+Performance+with+Heterogeneous+Memory+Tiering%3A+A+Case+Study+of+CXL-Based+Memory+Expansion
11. Tolerate It if You Cannot Reduce It: Handling Latency in Tiered Memory — approx. recent tiered-memory authors, 2024/2025
https://scholar.google.com/scholar?q=Tolerate+It+if+You+Cannot+Reduce+It%3A+Handling+Latency+in+Tiered+Memory
12. Can Hardware Outsmart Software in Tiered Memory Management? A CMM-H Case Study — approx. recent tiered-memory authors, 2024/2025
https://scholar.google.com/scholar?q=Can+Hardware+Outsmart+Software+in+Tiered+Memory+Management%3F+A+CMM-H+Case+Study
13. NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering — approx. recent CXL co-design authors, 2024/2025
https://scholar.google.com/scholar?q=NeoMem%3A+Hardware%2FSoftware+Co-Design+for+CXL-Native+Memory+Tiering
14. Survey of Disaggregated Memory: Cross-Layer Technique Insights for Next-Generation Datacenters — approx. survey authors, 2024/2025
https://scholar.google.com/scholar?q=Survey+of+Disaggregated+Memory%3A+Cross-Layer+Technique+Insights+for+Next-Generation+Datacenters
15. Disaggregated Memory in the Datacenter: A Survey — approx. survey authors, 2024/2025
https://scholar.google.com/scholar?q=Disaggregated+Memory+in+the+Datacenter%3A+A+Survey
16. AI Post Transformers: CXL Computational Memory Offloading for Lower Runtime — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-cxl-computational-memory-offloading-for-3b2124.mp3
17. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
18. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3
Interactive Visualization: Vistara Brings CXL Memory to Hyperscale
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof