AI Post Transformers

Weak-SIGReg for Stable Vision Transformer Training


Listen Later

This episode explores Weak-SIGReg, a lightweight covariance regularizer designed to prevent representation collapse in fragile supervised training, especially for small-data Vision Transformers. It explains how the method uses a sketched covariance matrix and an identity-matching penalty to keep hidden features decorrelated and similarly scaled at much lower cost than full covariance regularization. The discussion centers on CIFAR-100 results, where Weak-SIGReg dramatically improves a deliberately unstable ViT setup and also boosts a plain MLP, while offering little change on an already stable ResNet18. It also digs into the paper’s main caveat: the biggest gains appear when the baseline training recipe is badly broken, so the most interesting question is not just whether the regularizer works, but when it adds real value beyond simply fixing optimization and initialization.
Sources:
1. Weak-SIGReg: Covariance Regularization for Stable Deep Learning — Habibullah Akbar, 2026
http://arxiv.org/abs/2603.05924
2. Reducing Overfitting in Deep Networks by Decorrelating Representations — Michael Cogswell, Faruk Ahmed, Ross Girshick, Larry Zitnick, Dhruv Batra, 2015
https://scholar.google.com/scholar?q=Reducing+Overfitting+in+Deep+Networks+by+Decorrelating+Representations
3. Barlow Twins: Self-Supervised Learning via Redundancy Reduction — Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, Stephane Deny, 2021
https://scholar.google.com/scholar?q=Barlow+Twins%3A+Self-Supervised+Learning+via+Redundancy+Reduction
4. VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning — Adrien Bardes, Jean Ponce, Yann LeCun, 2021
https://scholar.google.com/scholar?q=VICReg%3A+Variance-Invariance-Covariance+Regularization+for+Self-Supervised+Learning
5. LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics — Randall Balestriero, Yann LeCun, 2025
https://scholar.google.com/scholar?q=LeJEPA%3A+Provable+and+Scalable+Self-Supervised+Learning+Without+the+Heuristics
6. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, et al., 2021
https://scholar.google.com/scholar?q=An+Image+is+Worth+16x16+Words%3A+Transformers+for+Image+Recognition+at+Scale
7. Training data-efficient image transformers & distillation through attention — Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, Herve Jegou, 2021
https://scholar.google.com/scholar?q=Training+data-efficient+image+transformers+%26+distillation+through+attention
8. How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers — Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, Lucas Beyer, 2021
https://scholar.google.com/scholar?q=How+to+train+your+ViT%3F+Data%2C+Augmentation%2C+and+Regularization+in+Vision+Transformers
9. Early Convolutions Help Transformers See Better — Tete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell, Piotr Dollar, Ross Girshick, 2021
https://scholar.google.com/scholar?q=Early+Convolutions+Help+Transformers+See+Better
10. Whitening for Self-Supervised Representation Learning — Aleksandr Ermolov, Aliaksandr Siarohin, Enver Sangineto, Nicu Sebe, 2020
https://arxiv.org/abs/2007.06346
11. Understanding Dimensional Collapse in Contrastive Self-Supervised Learning — Li Jing, Pascal Vincent, Yann LeCun, Yuandong Tian, 2021
https://arxiv.org/abs/2110.09348
12. UR-JEPA: Uniform Rectifiability as a Regularizer for Joint-Embedding Predictive Architectures — Triet M. Le, 2026
https://arxiv.org/abs/2606.01443
13. Stability of Transformers under Layer Normalization — Kelvin Kan, Xingjian Li, Benjamin J. Zhang, Tuhin Sahai, Stanley Osher, Krishna Kumar, Markos A. Katsoulakis, 2025
https://scholar.google.com/scholar?q=Stability+of+Transformers+under+Layer+Normalization
14. Transformers without Normalization — Jiachen Zhu, Xinlei Chen, Kaiming He, Yann LeCun, Zhuang Liu, 2025
https://scholar.google.com/scholar?q=Transformers+without+Normalization
15. Are Neurons Actually Collapsed? On the Fine-Grained Structure in Neural Representations — Yongyi Yang, Jacob Steinhardt, Wei Hu, 2023
https://scholar.google.com/scholar?q=Are+Neurons+Actually+Collapsed%3F+On+the+Fine-Grained+Structure+in+Neural+Representations
16. The Impact of Geometric Complexity on Neural Collapse in Transfer Learning — Michael Munn, Benoit Dherin, Javier Gonzalvo, 2024
https://scholar.google.com/scholar?q=The+Impact+of+Geometric+Complexity+on+Neural+Collapse+in+Transfer+Learning
17. AI Post Transformers: LeWorldModel: Stable Joint-Embedding World Models from Pixels — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-leworldmodel-stable-joint-embedding-worl-650f9f.mp3
18. AI Post Transformers: Muon Is Scalable for LLM Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-muon-is-scalable-for-llm-training-587ed8.mp3
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof