
Sign up to save your podcasts
Or


The piranha problem (too many large, independent effect sizes influence the same outcome) has received some attention on Andrew Gelman’s blog. But now it’s a paper! Chris Tosh (Memorial Sloan Kettering) talks about multiple views of the piranha problem and detecting the implausible scientific claims that are published. The butterfly effect makes an appearance.
If you enjoyed the science-vs-pseudoscience topics, you’ll enjoy this one.
0:00 - Coming up in the episode
2:35 - What is the Piranha Problem?
19:54 - Confusing effect sizes
23:11 - The "words & walking speed" study
26:22 - Declaration of independent variables
30:58 - Piranha theorems for correlations
37:07 - Piranha theorems for linear regression
40:37 - Piranha Theorems for mutual information
44:13 - Bounds on the independence of the covariates
46:12 - Applying the piranha theorem to real data
50:12 - Applying the piranha theorem across studies
54:05 - A Bayesian detour
1:00:12 - The butterfly effect & chaos
1:04:26 - Applying the piranha theorem to cancer research
Chris Holmes is Professor of Biostatistics at the University of Oxford and Programme Director for Health and Medical Sciences at The Alan Turing Institute. Chris’ research interests include Bayesian nonparametrics (which is the right kind of nonparametrics), statistical machine learning, genomics, and genetic epidemiology.
0:00 - Intro
Philosophy of Data Science Series
In the first keynote of the Philosophy of Data Science Series we have a 2-part interview with Deborah Mayo (Virginia Tech).
You can join our mail list at: https://www.podofasclepius.com/mail-list
We're always happy to hear your feedback and ideas - just post it in the YouTube comment section to start a conversation.
Thank you for your time and support of the series!
Topics:
0:00 - Preface to First Keynote Interview
Charlotte Deane | Bioinformatics, Deepmind's AlphaFold 2, and Llamas
Charlotte Deane (Oxford University) talks about statistical approaches to bioinformatics, the evolution of Google Deepmind's AlphaFold 2 & its place in protein informatics deep learning landscape. She also describes humanizing antibodies, and the increasing role of software engineers in statistical research groups. The topic of llamas, camels, and alpacas (and their unique place in proteomics research) makes a surprise visit.
[Note: This episode was originally published in January 2022, but the file contained a buffering error, which prevented the full interview from being played. This version, published Feb 1, 2022 contains the full interview.]
Charlotte Deane | Proteomics, AlphaFold 2, and Llamas
The philosophical community continuously aims to reconcile differing views on first person data and the consciousness of the mind. Is it possible to live without consciousness? Can one conceive thoughts without matching images to them? In this episode, Eric Schwitzgebel of the University of California tries to dissect such topics and questions to help us better understand the philosophical world.
Keywords: philosophy, epistemic data, first person data, stimulus error, imageless thought, consciousness
Starting a Statistics Consultancy | Janet Wittes
The following interview was a keynote fireside chat with Janet Wittes (Statistics Collaborative, Inc.) titled "Statisticians as Entrepreneurs". It was recorded for the BBSW 2021 Conference (Nov 3 - 5 in Foster City, CA).
References:
BBSW 2021 Conference: https://www.bbsw.org/bbsw2021
Topics:
0:00 Janet's background prior to founding Statistics Collaborative, Inc.
Jingyi Jessica Li | Advancing Statistical Genomics
Watch it on…. YouTube Podbean
Jingyi Jessica Li (UCLA) describes common statistical pitfalls in genomic data analysis & the statistical reasoning required to correct these mistakes.
Common themes throughout include:
Episode Topics
0:00 A major advancement in genomic data leads to new statistical techniques
2:15 Hypothesis-driven science & hypothesis-free data analysis
2:55 A ChIP Seq Example
8:00 Misformulation of sampling variability
16:55 A false analogy: the permutation test
19:03 Losing my p-value religion: the value of statistical packaging
24:30 The Clipper Framework for false discovery rate control
31:50 Non-parametric developments
37:55 Inferred covariates
46:00 PseudotimeDE: inferences of differential gene expression along cell pseudotime
47:10 Selective inference
49:25 What biological/physiological data will be incorporated in the future?
52:30 Statistics, computer science, data science, ML, biology
57:05 Machine learning and prediction
1:01:30 Sophisticated models vs sophisticated research
1:07:45 Peer review in science
1:13:05 Hypothesis-driven science vs cutting intellectual corners
1:18:12 What topic should the statistics community debate?
Mine Çetinkaya-Rundel | Advancing Open Access Data Science Education
Mine Çetinkaya-Rundel (Duke University) describes the current and future states of statistics and data science education. Then she discusses the process of building open access learning material.
0:00 - Introduction
Jingyi Jessica Li | Statistical Hypothesis Testing versus Machine Learning Binary Classification
Jingyi Jessica Li (UCLA) discusses her paper "Statistical Hypothesis Testing versus Machine Learning Binary Classification". Jingyi noticed several high-impact cancer research papers using multiple hypothesis testing for binary classification problems. Concerned that these papers had no guarantee on their claimed false discovery rates, Jingyi wrote a perspective article about clarifying hypothesis testing and binary classification to scientists.
#datascience #science #statistics
0:00 – Intro
From the publisher's feed