AI Post Transformers

Why LightGBM Made Boosted Trees Fast


Listen Later

This episode explores why LightGBM became a dominant tool for tabular machine learning by unpacking the algorithmic and systems ideas behind its speed. It explains how gradient boosting decision trees work, why split search becomes expensive on massive sparse datasets, and how LightGBM differs from neural-network-style training despite using gradient information. The discussion focuses on two core contributions: Gradient-based One-Side Sampling, which keeps high-gradient examples while subsampling easier ones without badly distorting split-gain estimates, and Exclusive Feature Bundling, which compresses sparse features by grouping columns that rarely activate together. Listeners would find it interesting for its clear account of how classical ideas like histograms, greedy tree growth, and graph coloring were combined into a highly practical system that reshaped real-world applications such as ranking, fraud detection, credit scoring, and forecasting.
Sources:
1. Why LightGBM Made Boosted Trees Fast
https://proceedings.neurips.cc/paper_files/paper/2017/file/6449f44a102fde848669bdd9eb6b76fa-Paper.pdf
2. Greedy Function Approximation: A Gradient Boosting Machine — Jerome H. Friedman, 2001
https://scholar.google.com/scholar?q=Greedy+Function+Approximation%3A+A+Gradient+Boosting+Machine
3. XGBoost: A Scalable Tree Boosting System — Tianqi Chen and Carlos Guestrin, 2016
https://scholar.google.com/scholar?q=XGBoost%3A+A+Scalable+Tree+Boosting+System
4. LightGBM: A Highly Efficient Gradient Boosting Decision Tree — Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, Tie-Yan Liu, 2017
https://scholar.google.com/scholar?q=LightGBM%3A+A+Highly+Efficient+Gradient+Boosting+Decision+Tree
5. CatBoost: Unbiased Boosting with Categorical Features — Liudmila Prokhorenkova, Gleb Gusev, Aleksandr Vorobev, Anna Veronika Dorogush, Andrey Gulin, 2018
https://scholar.google.com/scholar?q=CatBoost%3A+Unbiased+Boosting+with+Categorical+Features
6. Feature Hashing for Large Scale Multitask Learning — Kilian Weinberger, Anirban Dasgupta, Josh Attenberg, John Langford, Alex Smola, 2009
https://scholar.google.com/scholar?q=Feature+Hashing+for+Large+Scale+Multitask+Learning
7. An Upper Bound for the Chromatic Number of a Graph and Its Application to Timetabling Problems — D. J. A. Welsh and M. B. Powell, 1967
https://scholar.google.com/scholar?q=An+Upper+Bound+for+the+Chromatic+Number+of+a+Graph+and+Its+Application+to+Timetabling+Problems
8. New Methods to Color the Vertices of a Graph — Daniel Brélaz, 1979
https://scholar.google.com/scholar?q=New+Methods+to+Color+the+Vertices+of+a+Graph
9. Worst Case Behavior of Graph Coloring Algorithms — David S. Johnson, 1974
https://scholar.google.com/scholar?q=Worst+Case+Behavior+of+Graph+Coloring+Algorithms
10. A Communication-Efficient Parallel Algorithm for Decision Tree — Qi Meng, Guolin Ke, Taifeng Wang, Wei Chen, Qiwei Ye, Zhi-Ming Ma, Tie-Yan Liu, 2016
https://scholar.google.com/scholar?q=A+Communication-Efficient+Parallel+Algorithm+for+Decision+Tree
11. Stochastic Gradient Boosting — Jerome H. Friedman, 2002
https://scholar.google.com/scholar?q=Stochastic+Gradient+Boosting
12. Parallel Boosted Regression Trees for Web Search Ranking — Stephen Tyree, Kilian Q. Weinberger, Kunal Agrawal, and Jennifer Paykin, 2011
https://scholar.google.com/scholar?q=Parallel+Boosted+Regression+Trees+for+Web+Search+Ranking
13. Best-First Decision Tree Learning — Haijian Shi, 2007
https://scholar.google.com/scholar?q=Best-First+Decision+Tree+Learning
14. GPU-Acceleration for Large-Scale Tree Boosting — Huan Zhang, Si Si, and Cho-Jui Hsieh, 2017
https://scholar.google.com/scholar?q=GPU-Acceleration+for+Large-Scale+Tree+Boosting
15. Implementing machine learning methods with complex survey data: Lessons learned on the impacts of accounting sampling weights in gradient boosting — authors not identified in the provided snippet, recent (2020s)
https://scholar.google.com/scholar?q=Implementing+machine+learning+methods+with+complex+survey+data%3A+Lessons+learned+on+the+impacts+of+accounting+sampling+weights+in+gradient+boosting
16. Explainable boosting algorithms: sparse-group and interaction-aware variable selection in complex data — authors not identified in the provided snippet, recent (2020s)
https://scholar.google.com/scholar?q=Explainable+boosting+algorithms%3A+sparse-group+and+interaction-aware+variable+selection+in+complex+data
17. Multi-objective optimization of performance and interpretability of tabular supervised machine learning models — authors not identified in the provided snippet, recent (2020s)
https://scholar.google.com/scholar?q=Multi-objective+optimization+of+performance+and+interpretability+of+tabular+supervised+machine+learning+models
18. AI Post Transformers: Breiman's Two Cultures of Statistical Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-breimans-two-cultures-of-statistical-mod-71e49f.mp3
Interactive Visualization: Why LightGBM Made Boosted Trees Fast
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof