Data Science Decoded

Data Science #15 - The First Decision Tree Algorithm (1963)


Listen Later

In the 15th episode we went over the paper "Problems in the Analysis of Survey Data, and a Proposal" by James N. Morgan and John A. Sonquist from 1964.

It highlights seven key issues in analyzing complex survey data, such as high dimensionality, categorical variables, measurement errors, sample variability, intercorrelations, interaction effects, and causal chains.

These challenges complicate efforts to draw meaningful conclusions about relationships between factors like income, education, and occupation.

To address these problems, the authors propose a method that sequentially splits data by identifying features that reduce unexplained variance, much like modern decision trees.

The method focuses on maximizing explained variance (SSE), capturing interaction effects, and accounting for sample variability.

It handles both categorical and continuous variables while respecting logical causal priorities.

This paper has had a significant influence on modern data science and AI, laying the groundwork for decision trees, CART, random forests, and boosting algorithms. Its method of splitting data to reduce error, handle interactions, and respect feature hierarchies is foundational in many machine learning models used today.

...more
View all episodesView all episodes
Download on the App Store

Data Science DecodedBy Mike E

  • 3.8
  • 3.8
  • 3.8
  • 3.8
  • 3.8

3.8

5 ratings


More shows like Data Science Decoded

View all
Freakonomics Radio by Freakonomics Radio + Stitcher

Freakonomics Radio

32,020 Listeners

The Joe Rogan Experience by Joe Rogan

The Joe Rogan Experience

229,278 Listeners

StarTalk Radio by Neil deGrasse Tyson

StarTalk Radio

14,322 Listeners

More or Less by BBC Radio 4

More or Less

887 Listeners

Talk Python To Me by Michael Kennedy

Talk Python To Me

585 Listeners

The Quanta Podcast by Quanta Magazine

The Quanta Podcast

529 Listeners

Science Friday by Science Friday and WNYC Studios

Science Friday

6,420 Listeners

Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

Super Data Science: ML & AI Podcast with Jon Krohn

302 Listeners

Pod Save America by Crooked Media

Pod Save America

87,414 Listeners

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas by Sean Carroll | Wondery

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas

4,178 Listeners

Practical AI by Practical AI LLC

Practical AI

210 Listeners

Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

95 Listeners

Hard Fork by The New York Times

Hard Fork

5,526 Listeners

The Rest Is History by Goalhanger

The Rest Is History

15,222 Listeners

The Astrophysics Podcast by Paul Duffell

The Astrophysics Podcast

54 Listeners