AI Post Transformers

Fast FPGA BDT Inference for LHC Triggers


Listen Later

This episode explores how boosted decision trees can be compiled directly into FPGA firmware for ultra-low-latency particle-physics triggers at the Large Hadron Collider. It explains why this setting favors shallow, quantized tree ensembles over larger neural networks: trigger decisions must happen within a tiny hardware budget, with strict limits on latency, power, and on-chip resources. The discussion focuses on a concrete benchmark where a 100-tree, depth-4 gradient-boosted model for five-class jet tagging is mapped to a Xilinx VU9P FPGA and compared against a similarly deployed multilayer perceptron. Listeners would find it interesting because it shows how model choice changes when every nanosecond matters, and how familiar ML methods can become hardwired decision circuits rather than conventional software inference.
Sources:
1. Fast inference of Boosted Decision Trees in FPGAs for particle physics — Sioni Summers, Giuseppe Di Guglielmo, Javier Duarte, Philip Harris, Duc Hoang, Sergo Jindariani, Edward Kreinar, Vladimir Loncar, Jennifer Ngadiuba, Maurizio Pierini, Dylan Rankin, Nhan Tran, Zhenbin Wu, 2020
http://arxiv.org/abs/2002.02534
2. Greedy Function Approximation: A Gradient Boosting Machine — Jerome H. Friedman, 2001
https://scholar.google.com/scholar?q=Greedy+Function+Approximation%3A+A+Gradient+Boosting+Machine
3. XGBoost: A Scalable Tree Boosting System — Tianqi Chen, Carlos Guestrin, 2016
https://scholar.google.com/scholar?q=XGBoost%3A+A+Scalable+Tree+Boosting+System
4. LightGBM: A Highly Efficient Gradient Boosting Decision Tree — Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, Tie-Yan Liu, 2017
https://scholar.google.com/scholar?q=LightGBM%3A+A+Highly+Efficient+Gradient+Boosting+Decision+Tree
5. CatBoost: unbiased boosting with categorical features — Liudmila Prokhorenkova, Gleb Gusev, Aleksandr Vorobev, Anna Veronika Dorogush, Andrey Gulin, 2018
https://scholar.google.com/scholar?q=CatBoost%3A+unbiased+boosting+with+categorical+features
6. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference — Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, Dmitry Kalenichenko, 2018
https://scholar.google.com/scholar?q=Quantization+and+Training+of+Neural+Networks+for+Efficient+Integer-Arithmetic-Only+Inference
7. Quantizing deep convolutional networks for efficient inference: A whitepaper — Raghuraman Krishnamoorthi, 2018
https://scholar.google.com/scholar?q=Quantizing+deep+convolutional+networks+for+efficient+inference%3A+A+whitepaper
8. Post-training 4-bit quantization of convolution networks for rapid-deployment — Ron Banner, Yury Nahshan, Elad Hoffer, Daniel Soudry, 2019
https://scholar.google.com/scholar?q=Post-training+4-bit+quantization+of+convolution+networks+for+rapid-deployment
9. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers — Elias Frantar, Saleh Ashkboos, Torsten Hoefler, Dan Alistarh, 2022
https://scholar.google.com/scholar?q=GPTQ%3A+Accurate+Post-Training+Quantization+for+Generative+Pre-trained+Transformers
10. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, Song Han, 2023
https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration
11. Fast inference of deep neural networks in FPGAs for particle physics — Javier Duarte, Song Han, Philip Harris, Sergo Jindariani, Edward Kreinar, Benjamin Kreis, Jennifer Ngadiuba, Maurizio Pierini, Ryan Rivera, Nhan Tran, Zhenbin Wu, 2018
https://scholar.google.com/scholar?q=Fast+inference+of+deep+neural+networks+in+FPGAs+for+particle+physics
12. Efficient, reliable and fast high-level triggering using a bonsai boosted decision tree — V. V. Gligorov, M. Williams, 2013
https://scholar.google.com/scholar?q=Efficient%2C+reliable+and+fast+high-level+triggering+using+a+bonsai+boosted+decision+tree
13. Boosted Decision Trees in the Level-1 Muon Endcap Trigger at CMS — CMS Collaboration, 2018
https://scholar.google.com/scholar?q=Boosted+Decision+Trees+in+the+Level-1+Muon+Endcap+Trigger+at+CMS
14. Scalable inference of decision tree ensembles: Flexible design for CPU-FPGA platforms — Muhsen Owaida, Hantian Zhang, Ce Zhang, Gustavo Alonso, 2017
https://scholar.google.com/scholar?q=Scalable+inference+of+decision+tree+ensembles%3A+Flexible+design+for+CPU-FPGA+platforms
15. Machine learning at the energy and intensity frontiers of particle physics — A. Radovic et al., 2018
https://scholar.google.com/scholar?q=Machine+learning+at+the+energy+and+intensity+frontiers+of+particle+physics
16. Low latency transformer inference on FPGAs for physics applications with hls4ml — Zhixing Jiang et al., 2025
https://scholar.google.com/scholar?q=Low+latency+transformer+inference+on+FPGAs+for+physics+applications+with+hls4ml
17. Ultrafast jet classification at the HL-LHC — Patrick Odagiu et al., 2024
https://scholar.google.com/scholar?q=Ultrafast+jet+classification+at+the+HL-LHC
18. AI Post Transformers: FlatAttention for Tile-Based Accelerator Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-flatattention-for-tile-based-accelerator-56e6ca.mp3
19. AI Post Transformers: Computation-Bandwidth-Memory Trade-offs for AI Infrastructure — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-computation-bandwidth-memory-trade-offs-a83f2b.mp3
20. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3
Interactive Visualization: Fast FPGA BDT Inference for LHC Triggers
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof