This episode explores Paul Werbos’s 2004 review of reverse differentiation and argues that reverse-mode automatic differentiation, backpropagation, hand-coded adjoints, and adjoint circuits are largely the same core idea expressed in different technical communities. It explains the mechanics of automatic differentiation and reverse mode in clear terms, then traces how these methods diverged historically and why that fragmentation slowed progress in fields like neural networks, control, and scientific computing. The discussion highlights Werbos’s main claim that better integrated, derivative-aware software could make advanced nonlinear modeling and intelligent control far more practical, while also questioning how much evidence supports that agenda beyond synthesis and historical interpretation. Listeners would find it interesting for its sharp distinction between gradients as infrastructure versus models or optimizers, and for its perspective on how today’s differentiable programming ecosystem was once a contested software vision.
Sources:
1. Reverse-Mode Differentiation Across AD and Neural Nets
https://www.werbos.com/AD2004.pdf
2. A Simple Automatic Derivative Evaluation Program — R. E. Wengert, 1964
https://scholar.google.com/scholar?q=A+Simple+Automatic+Derivative+Evaluation+Program
3. Taylor Expansion of the Accumulated Rounding Error — Seppo Linnainmaa, 1976
https://scholar.google.com/scholar?q=Taylor+Expansion+of+the+Accumulated+Rounding+Error
4. The Complexity of Partial Derivatives — Walter Baur and Volker Strassen, 1983
https://scholar.google.com/scholar?q=The+Complexity+of+Partial+Derivatives
5. Automatic Differentiation in Machine Learning: a Survey — Atılım Güneş Baydin, Barak A. Pearlmutter, Alexey Andreyevich Radul, Jeffrey Mark Siskind, 2018
https://scholar.google.com/scholar?q=Automatic+Differentiation+in+Machine+Learning%3A+a+Survey
6. Learning Representations by Back-Propagating Errors — David E. Rumelhart, Geoffrey E. Hinton, Ronald J. Williams, 1986
https://scholar.google.com/scholar?q=Learning+Representations+by+Back-Propagating+Errors
7. Backpropagation Through Time: What It Does and How to Do It — Paul J. Werbos, 1990
https://scholar.google.com/scholar?q=Backpropagation+Through+Time%3A+What+It+Does+and+How+to+Do+It
8. Backpropagation Applied to Handwritten Zip Code Recognition — Yann LeCun, Bernhard Boser, John S. Denker, Don Henderson, Richard E. Howard, Wayne Hubbard, Lawrence D. Jackel, 1989
https://scholar.google.com/scholar?q=Backpropagation+Applied+to+Handwritten+Zip+Code+Recognition
9. Gradient-Based Learning Applied to Document Recognition — Yann LeCun, Léon Bottou, Yoshua Bengio, Patrick Haffner, 1998
https://scholar.google.com/scholar?q=Gradient-Based+Learning+Applied+to+Document+Recognition
10. Neuro-Dynamic Programming: An Overview — Dimitri P. Bertsekas, John N. Tsitsiklis, 1995
https://scholar.google.com/scholar?q=Neuro-Dynamic+Programming%3A+An+Overview
11. Neuro-Dynamic Programming — Dimitri P. Bertsekas, John N. Tsitsiklis, 1996
https://scholar.google.com/scholar?q=Neuro-Dynamic+Programming
12. Approximate Dynamic Programming and Reinforcement Learning — Lucian Bușoniu, Bart De Schutter, Robert Babuška, 2010
https://scholar.google.com/scholar?q=Approximate+Dynamic+Programming+and+Reinforcement+Learning
13. An Approximate Dynamic Programming Algorithm for Large-Scale Fleet Management: A Case Application — Hugo P. Simão, Jeff Day, Abraham P. George, Ted Gifford, John Nienow, Warren B. Powell, 2009
https://scholar.google.com/scholar?q=An+Approximate+Dynamic+Programming+Algorithm+for+Large-Scale+Fleet+Management%3A+A+Case+Application
14. Some New Tools for Prediction and Analysis in the Behavioral Sciences — Paul J. Werbos, 1974
https://scholar.google.com/scholar?q=Some+New+Tools+for+Prediction+and+Analysis+in+the+Behavioral+Sciences
15. The Difficulty of Learning Long-Term Dependencies with Gradient Descent is Officially Overcome — Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, Jürgen Schmidhuber, 2001
https://scholar.google.com/scholar?q=The+Difficulty+of+Learning+Long-Term+Dependencies+with+Gradient+Descent+is+Officially+Overcome
16. Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation — Andreas Griewank, 2000
https://scholar.google.com/scholar?q=Evaluating+Derivatives%3A+Principles+and+Techniques+of+Algorithmic+Differentiation
17. Differential Dynamic Programming — David H. Jacobson, David Q. Mayne, 1970
https://scholar.google.com/scholar?q=Differential+Dynamic+Programming
18. Backpropagation-free training of deep physical neural networks — Ali Momeni, Babak Rahmani, Matthieu Mallejac, Philipp Del Hougne, Romain Fleury, 2023
https://scholar.google.com/scholar?q=Backpropagation-free+training+of+deep+physical+neural+networks
19. Fully forward mode training for optical neural networks — Zhiwei Xue, Tiankuang Zhou, Zhihao Xu, Shaoliang Yu, Qionghai Dai, Lu Fang, 2024
https://scholar.google.com/scholar?q=Fully+forward+mode+training+for+optical+neural+networks
20. Brain-like training of a pre-sensor optical neural network with a backpropagation-free algorithm — Zheng Huang, Conghe Wang, Caihua Zhang, Wanxin Shi, Shukai Wu, Sigang Yang, Hongwei Chen, 2025
https://scholar.google.com/scholar?q=Brain-like+training+of+a+pre-sensor+optical+neural+network+with+a+backpropagation-free+algorithm
21. Backpropagation-Free Deep Learning with Recursive Local Representation Alignment — Alexander G. Ororbia, Ankur Mali, Daniel Kifer, C. Lee Giles, 2023
https://scholar.google.com/scholar?q=Backpropagation-Free+Deep+Learning+with+Recursive+Local+Representation+Alignment
22. Exploring the Promise and Limits of Real-Time Recurrent Learning — Kazuki Irie, Anand Gopalakrishnan, Jürgen Schmidhuber, 2023
https://scholar.google.com/scholar?q=Exploring+the+Promise+and+Limits+of+Real-Time+Recurrent+Learning
23. Real-Time Recurrent Reinforcement Learning — Julian Lemmel, Radu Grosu, 2023/2025
https://scholar.google.com/scholar?q=Real-Time+Recurrent+Reinforcement+Learning
24. Second-order forward-mode optimization of recurrent neural networks for neuroscience — Youjing Yu, Rui Xia, Qingxi Ma, Máté Lengyel, Guillaume Hennequin, 2024
https://scholar.google.com/scholar?q=Second-order+forward-mode+optimization+of+recurrent+neural+networks+for+neuroscience
25. Dynamic predictive coding: A model of hierarchical sequence learning and prediction in the neocortex — Linxing Preston Jiang, Rajesh P. N. Rao, 2024
https://scholar.google.com/scholar?q=Dynamic+predictive+coding%3A+A+model+of+hierarchical+sequence+learning+and+prediction+in+the+neocortex
26. Predictive coding networks for temporal prediction — Beren Millidge, Mufeng Tang, Mahyar Osanlouy, Nicol S. Harper, Rafal Bogacz, 2024
https://scholar.google.com/scholar?q=Predictive+coding+networks+for+temporal+prediction
27. Where is the error? Hierarchical predictive coding through dendritic error computation — Fabian A. Mikulasch, Lucas Rudelt, Michael Wibral, Viola Priesemann, 2023
https://scholar.google.com/scholar?q=Where+is+the+error%3F+Hierarchical+predictive+coding+through+dendritic+error+computation
28. AI Post Transformers: Long Short-Term Memory and Vanishing Gradients — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-long-short-term-memory-and-vanishing-gra-72448c.mp3
29. AI Post Transformers: When Spectral Gradient Updates Help Deep Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-when-spectral-gradient-updates-help-deep-9c8441.mp3
30. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3
Interactive Visualization: Reverse-Mode Differentiation Across AD and Neural Nets