This episode explores EventTensor, a compiler abstraction from Carnegie Mellon and collaborators (presented at MLSys 2026) that treats synchronization events as first-class tensors for compiling GPU megakernels. The discussion covers how encoding true data dependencies—rather than waiting for entire kernels to finish—enables fine-grained scheduling, illustrated through a split-K summation example and a symbolic batch-size template that avoids recompilation when shapes change. A key focus is how the system handles Mixture-of-Experts routing, where dependencies aren't known until runtime, via data-dependent event counters and task triggering computed from router outputs. The hosts also unpack the tradeoffs between static and dynamic scheduling, showing that static wins on predictable dense workloads while dynamic pays off only under genuine irregularity like MoE. Benchmark results show up to 1.40x speedups over cuBLAS+NCCL, 1.23x over Triton/FlashInfer on MoE layers, and end-to-end gains of 1.48x over vLLM, making this a concrete look at how compile-time and runtime scheduling can be unified without sacrificing performance.
Sources:
1. Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel — Hongyi Jin, Bohan Hou, Guanjie Wang, Ruihang Lai, Jinqi Chen, Zihao Ye, Yaxing Cai, Yixin Dong, Xinhao Cheng, Zhihao Zhang, Yilong Zhao, Yingyi Huang, Lijie Yang, Jinchen Jiang, Gabriele Oliaro, Jianan Ji, Xupeng Miao, Vinod Grover, Todd C. Mowry, Zhihao Jia, Tianqi Chen, 2026
http://arxiv.org/abs/2604.13327
2. Legion: Expressing Locality and Independence with Logical Regions — Michael Bauer, Sean Treichler, Elliott Slaughter, Alex Aiken, 2012
https://scholar.google.com/scholar?q=Legion%3A+Expressing+Locality+and+Independence+with+Logical+Regions
3. StarPU: A Unified Platform for Task Scheduling on Heterogeneous Multicore Architectures — Cédric Augonnet, Samuel Thibault, Raymond Namyst, Pierre-André Wacrenier, 2011
https://scholar.google.com/scholar?q=StarPU%3A+A+Unified+Platform+for+Task+Scheduling+on+Heterogeneous+Multicore+Architectures
4. Dynamic Control Flow in Large-Scale Machine Learning — Yuan Yu, Martín Abadi, Paul Barham, Eugene Brevdo, Mike Burrows, Andy Davis, Jeff Dean, Sanjay Ghemawat, Tim Harley, Peter Hawkins, Mark Hong, Rajat Monga, Derek Murray, Xiaoqiang Zheng, and others (Google Brain), 2018
https://scholar.google.com/scholar?q=Dynamic+Control+Flow+in+Large-Scale+Machine+Learning
5. Ray: A Distributed Framework for Emerging AI Applications — Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, Ion Stoica, 2018
https://scholar.google.com/scholar?q=Ray%3A+A+Distributed+Framework+for+Emerging+AI+Applications
6. Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs — Cheng, X., Zhang, Z., Zhou, Y., Ji, J., Jiang, J., Zhao, Z., et al. (overlapping author list with this paper), 2025
https://scholar.google.com/scholar?q=Mirage+Persistent+Kernel%3A+A+Compiler+and+Runtime+for+Mega-Kernelizing+Tensor+Programs
7. Look ma, no bubbles! Designing a low-latency megakernel for Llama-1B — Spector, B., Juravsky, J., Sul, S., Dugan, O., Lim, D., Fu, D., Arora, S., R, C., 2025
https://scholar.google.com/scholar?q=Look+ma%2C+no+bubbles%21+Designing+a+low-latency+megakernel+for+Llama-1B
8. A Framework for Fine-Grained Synchronization of Dependent GPU Kernels (CuSync) — Jangda, A., Maleki, S., Dehnavi, M. M., Musuvathi, M., Saarikivi, O., 2024
https://scholar.google.com/scholar?q=A+Framework+for+Fine-Grained+Synchronization+of+Dependent+GPU+Kernels+%28CuSync%29
9. Graphene: An IR for Optimized Tensor Computations on GPUs — Hagedorn, B., Fan, B., Chen, H., Cecka, C., Garland, M., Grover, V., 2023
https://scholar.google.com/scholar?q=Graphene%3A+An+IR+for+Optimized+Tensor+Computations+on+GPUs
10. FlashMoE: Fast Distributed MoE in a Single Kernel — Aimuyo, O. J., Oh, B., Singh, R., 2025
https://scholar.google.com/scholar?q=FlashMoE%3A+Fast+Distributed+MoE+in+a+Single+Kernel
Interactive Visualization: Prompt Boundary-Aware Scheduling with Event Tensors for Dynamic Kernels