AI Post Transformers

Learning in Random Nets and Generalization


Listen Later

This episode explores Marvin Minsky’s 1961 paper on whether mostly random neural networks can learn useful behavior simply by reinforcing successful responses. It explains how the paper distinguishes rote memory, associative recall, pattern recognition, and true generalization, arguing that reward signals alone are not enough unless the system already has a meaningful notion of similarity between situations. The discussion places that idea in context with early machine learning work like Rosenblatt’s perceptron and Samuel’s checkers program, then connects it to later, more disciplined descendants such as echo state networks and random features. Listeners get a sharp historical view of a debate that still matters now: whether intelligence comes from discovering good representations or from selecting among structures that were already there.
Sources:
1. Learning in Random Nets and Generalization
https://stacks.stanford.edu/file/druid:yr384hg3073/yr384hg3073.pdf
2. Learning in Random Nets — Marvin Minsky, Oliver G. Selfridge, 1961
https://scholar.google.com/scholar?q=Learning+in+Random+Nets
3. The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain — Frank Rosenblatt, 1958
https://scholar.google.com/scholar?q=The+Perceptron%3A+A+Probabilistic+Model+for+Information+Storage+and+Organization+in+the+Brain
4. The Echo State Approach to Analysing and Training Recurrent Neural Networks — Herbert Jaeger, 2001
https://scholar.google.com/scholar?q=The+Echo+State+Approach+to+Analysing+and+Training+Recurrent+Neural+Networks
5. Random Features for Large-Scale Kernel Machines — Ali Rahimi, Benjamin Recht, 2007
https://scholar.google.com/scholar?q=Random+Features+for+Large-Scale+Kernel+Machines
6. The Use of Multiple Measurements in Taxonomic Problems — R. A. Fisher, 1936
https://scholar.google.com/scholar?q=The+Use+of+Multiple+Measurements+in+Taxonomic+Problems
7. Nearest Neighbor Pattern Classification — Thomas M. Cover, Peter E. Hart, 1967
https://scholar.google.com/scholar?q=Nearest+Neighbor+Pattern+Classification
8. Support-Vector Networks — Corinna Cortes, Vladimir Vapnik, 1995
https://scholar.google.com/scholar?q=Support-Vector+Networks
9. Gradient-Based Learning Applied to Document Recognition — Yann LeCun, Leon Bottou, Yoshua Bengio, Patrick Haffner, 1998
https://scholar.google.com/scholar?q=Gradient-Based+Learning+Applied+to+Document+Recognition
10. Some Studies in Machine Learning Using the Game of Checkers — Arthur L. Samuel, 1959
https://scholar.google.com/scholar?q=Some+Studies+in+Machine+Learning+Using+the+Game+of+Checkers
11. Generalization of Pattern Recognition in a Self-Organizing System — B. G. Farley and W. A. Clark, 1954
https://scholar.google.com/scholar?q=Generalization+of+Pattern+Recognition+in+a+Self-Organizing+System
12. A Heterarchy of Values Determined by the Topology of Nervous Nets — Warren S. McCulloch, 1945
https://scholar.google.com/scholar?q=A+Heterarchy+of+Values+Determined+by+the+Topology+of+Nervous+Nets
13. Asymptotics of Random Feature Regression Beyond the Linear Scaling Regime — Hong Hu, Yue M. Lu, Theodor Misiakiewicz, 2024
https://scholar.google.com/scholar?q=Asymptotics+of+Random+Feature+Regression+Beyond+the+Linear+Scaling+Regime
14. Power-Law Spectrum of the Random Feature Model — Elliot Paquette, Ke Liang Xiao, Yizhe Zhu, 2026
https://scholar.google.com/scholar?q=Power-Law+Spectrum+of+the+Random+Feature+Model
15. Local to Global: Learning Dynamics and Effect of Initialization for Transformers — Ashok Vardhan Makkuva, Marco Bondaschi, Chanakya Ekbote, Adway Girish, Alliot Nagle, Hyeji Kim, Michael Gastpar, 2024
https://scholar.google.com/scholar?q=Local+to+Global%3A+Learning+Dynamics+and+Effect+of+Initialization+for+Transformers
16. Augmenting Language Models with Long-Term Memory — Weizhi Wang, Li Dong, Hao Cheng, Xiaodong Liu, Xifeng Yan, Jianfeng Gao, Furu Wei, 2023
https://scholar.google.com/scholar?q=Augmenting+Language+Models+with+Long-Term+Memory
17. MemoryBank: Enhancing Large Language Models with Long-Term Memory — Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, Yanlin Wang, 2023
https://scholar.google.com/scholar?q=MemoryBank%3A+Enhancing+Large+Language+Models+with+Long-Term+Memory
18. HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models — Bernal Jiménez Gutiérrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, Yu Su, 2024
https://scholar.google.com/scholar?q=HippoRAG%3A+Neurobiologically+Inspired+Long-Term+Memory+for+Large+Language+Models
19. The Learnability of In-Context Learning — Noam Wies, Yoav Levine, Amnon Shashua, 2023
https://scholar.google.com/scholar?q=The+Learnability+of+In-Context+Learning
20. Understanding In-Context Learning in Transformers and LLMs by Learning to Learn Discrete Functions — Satwik Bhattamishra, Arkil Patel, Phil Blunsom, Varun Kanade, 2023
https://scholar.google.com/scholar?q=Understanding+In-Context+Learning+in+Transformers+and+LLMs+by+Learning+to+Learn+Discrete+Functions
21. Learning without Training: The Implicit Dynamics of In-Context Learning — Benoit Dherin, Michael Munn, Hanna Mazzawi, Michael Wunder, Javier Gonzalvo, 2025
https://scholar.google.com/scholar?q=Learning+without+Training%3A+The+Implicit+Dynamics+of+In-Context+Learning
22. AI Post Transformers: Long Short-Term Memory and Vanishing Gradients — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-long-short-term-memory-and-vanishing-gra-72448c.mp3
23. AI Post Transformers: Latent Space as a New Computational Paradigm — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-latent-space-as-a-new-computational-para-810f39.mp3
24. AI Post Transformers: When Spectral Gradient Updates Help Deep Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-when-spectral-gradient-updates-help-deep-9c8441.mp3
25. AI Post Transformers: LeWorldModel: Stable Joint-Embedding World Models from Pixels — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-leworldmodel-stable-joint-embedding-worl-650f9f.mp3
Interactive Visualization: Learning in Random Nets and Generalization
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof