Summary
What separates a valid A/B test from an expensive guess? Priya Singhee, VP of Enterprise Analytics & Data Science at Early Warning — the bank-owned consortium that fights payment fraud and operates Zelle, which processed a trillion dollars last year — joins host Ashley Stirrup to share the experimentation playbook she built leading storefront analytics at Wayfair. She walks through the four questions to ask before launching any A/B test, why 85 to 90% of tests are supposed to fail, how pre-registration and kill criteria stop p-hacking before it starts, the pitfalls that fake wins (novelty effects, hidden heterogeneity, multiple comparisons), and how to roll out winners with gradual ramps and long-running holdouts. A practical episode for product managers, engineers, data scientists, and growth leaders building rigorous experimentation programs.
Chapters
00:45 Meet Early Warning: fraud detection, Zelle, and a trillion dollars in payments
02:00 Wayfair and optimizing every step of the storefront funnel
03:05 The four questions to ask before any A/B test
05:15 Test setup best practices: hypotheses, guardrails, power, and pre-registration
07:40 Why 85 to 90% of tests fail and why that's a learning agenda
09:25 Novelty effects, hidden heterogeneity, and the multiple comparisons problem
12:55 Pre-registration, kill criteria, and stopping p-hacking
15:10 Rolling out winners: gradual ramps and long-running holdouts
17:15 Causal inference when you can't A/B test
20:05 The case for more A/B testing, not less
Takeaways
- Run the four-question framework before any test: clean randomization, a plausible effect size for your traffic, a reversible and cheap change, and a falsifiable hypothesis.
- Treat A/B testing as a learning agenda: 85 to 90% of tests are supposed to fail, and a suspiciously high win rate is a red flag, not a trophy.
- Pre-register the full analysis plan, including hypothesis, mechanism, primary metric, exact statistical test, and subgroups, so p-hacking can't creep in when a test goes sideways.
- Define kill criteria and success, failure, and guardrail-dip actions before launch, with leadership sign-off, so nobody chases a loss into a fake win.
- Log every test and its learnings where the whole organization can see them; that reinforcement loop is what separates world-class experimentation programs.
Connect with the Guest
Priya Singhee LinkedIn: https://www.linkedin.com/in/priya-singhee/
About the guest: https://www.growthbook.io/podcast/guests/priya-singhee
Episode page: https://www.growthbook.io/podcast/episode/1-37
Company Website: https://www.earlywarning.com
Sponsor
GrowthBook is the warehouse-native platform for experimentation, feature flags, and product analytics trusted by AI-native product teams at 3,000+ companies worldwide.
Go to http://growthbook.io