
Sign up to save your podcasts
Or


[00:00] Intro: Research on multi-task learning in LLMs
[00:38] Balancing safety and performance in multilingual settings
[01:17] Model merging techniques explored
[02:16] Model merging outperforms data mixing
[02:49] Merging monolingual models improves multilingual capabilities
[03:28] Key ablation studies
[04:14] Safety and performance evaluation metrics
[04:52] Effectiveness variations across languages
[05:26] Safety model weight impact in linear merging
[06:00] Insights on merging and preference training
[06:25] Comparison to existing research
[06:56] Implications for LLM development
[07:29] Limitations and future research
Authors: Aakanksha, Arash Ahmadian, Seraphina Goldfarb-Tarrant, Beyza Ermis, Marzieh Fadaee, Sara Hooker
Affiliations: Cohere For AI, Cohere
Abstract: Large Language Models (LLMs) have been adopted and deployed worldwide for a broad variety of applications. However, ensuring their safe use remains a significant challenge. Preference training and safety measures often overfit to harms prevalent in Western-centric datasets, and safety protocols frequently fail to extend to multilingual settings. In this work, we explore model merging in a diverse multi-task setting, combining safety and general-purpose tasks within a multilingual context. Each language introduces unique and varied learning challenges across tasks. We find that objective-based merging is more effective than mixing data, with improvements of up to 8% and 10% in general performance and safety respectively. We also find that language-based merging is highly effective -- by merging monolingually fine-tuned models, we achieve a 4% increase in general performance and 7% reduction in harm across all languages on top of the data mixtures method using the same available data. Overall, our comprehensive study of merging approaches provides a useful framework for building strong and safe multilingual models.
Link: https://arxiv.org/abs/2410.10801
[00:00] Intro: CoTracker 3 paper
[00:17] Challenge: Point tracker training
[00:51] Existing methods' limitations
[01:05] CoTracker 3 innovations
[01:41] Novel semi-supervised training
[02:24] Online vs offline versions
[03:02] Performance improvements
[03:47] Ablation study insights
[04:29] Limitations and future work
Authors: Nikita Karaev, Iurii Makarov, Jianyuan Wang, Natalia Neverova, Andrea Vedaldi, Christian Rupprecht
Affiliations: Meta AI; Visual Geometry Group, University of Oxford
Abstract: Most state-of-the-art point trackers are trained on synthetic data due to the difficulty of annotating real videos for this task. However, this can result in suboptimal performance due to the statistical gap between synthetic and real videos. In order to understand these issues better, we introduce CoTracker3, comprising a new tracking model and a new semi-supervised training recipe. This allows real videos without annotations to be used during training by generating pseudo-labels using off-the-shelf teachers. The new model eliminates or simplifies components from previous trackers, resulting in a simpler and often smaller architecture. This training scheme is much simpler than prior work and achieves better results using 1,000 times less data. We further study the scaling behaviour to understand the impact of using more real unsupervised data in point tracking. The model is available in online and offline variants and reliably tracks visible and occluded points.
Link: https://arxiv.org/abs/2410.11831
[00:00] Introduction
[00:30] Consistent unit norm normalization in NGPT
[01:08] Mathematical mechanism behind faster convergence
[01:52] Elimination of weight decay in NGPT
[02:21] Role of learnable eigen learning rates in optimization
[03:04] Discussion on training speedup vs. per-step computation time
[03:46] Condition number differences between GPT and NGPT
[04:18] Ablation studies on scaling factors
[04:53] NGPT's relationship to Riemannian optimization
[05:27] Future research
[06:02] Takeaways for practitioners
Authors: Ilya Loshchilov, Cheng-Ping Hsieh, Simeng Sun, Boris Ginsburg
Affiliations: NVIDIA
Abstract: We propose a novel neural network architecture, the normalized Transformer (nGPT) with representation learning on the hypersphere. In nGPT, all vectors forming the embeddings, MLP, attention matrices and hidden states are unit norm normalized. The input stream of tokens travels on the surface of a hypersphere, with each layer contributing a displacement towards the target output predictions. These displacements are defined by the MLP and attention blocks, whose vector components also reside on the same hypersphere. Experiments show that nGPT learns much faster, reducing the number of training steps required to achieve the same accuracy by a factor of 4 to 20, depending on the sequence length.
[00:00] Intro
Authors: Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan, Huayu Chen, Zhengyi Wang, Ke Xu, Hang Su, Jun Zhu
Affiliations: Tsinghua University
Abstract: Bimanual manipulation is essential in robotics, yet developing foundation models is extremely challenging due to the inherent complexity of coordinating two robot arms (leading to multi-modal action distributions) and the scarcity of training data. In this paper, we present the Robotics Diffusion Transformer (RDT), a pioneering diffusion foundation model for bimanual manipulation. RDT builds on diffusion models to effectively represent multi-modality, with innovative designs of a scalable Transformer to deal with the heterogeneity of multi-modal inputs and to capture the nonlinearity and high frequency of robotic data. To address data scarcity, we further introduce a Physically Interpretable Unified Action Space, which can unify the action representations of various robots while preserving the physical meanings of original actions, facilitating learning transferrable physical knowledge. With these designs, we managed to pre-train RDT on the largest collection of multi-robot datasets to date and scaled it up to 1.2B parameters, which is the largest diffusion-based foundation model for robotic manipulation. We finally fine-tuned RDT on a self-created multi-task bimanual dataset with over 6K+ episodes to refine its manipulation capabilities. Experiments on real robots demonstrate that RDT significantly outperforms existing methods. It exhibits zero-shot generalization to unseen objects and scenes, understands and follows language instructions, learns new skills with just 1~5 demonstrations, and effectively handles complex, dexterous tasks. We refer to this https URL for the code and videos.
Link: https://rdt-robotics.github.io/rdt-robotics/
[00:00] Continuous-time consistency models (CTMs)
[00:20] Limitations of existing CTMs
[00:51] TrigFlow: New CTM formulation
[01:42] CTM training instability
[02:21] Training objective modifications
[02:55] Scaling CTMs to 1.5 billion parameters
[03:37] Comparison with state-of-the-art models
[04:14] Consistency training vs. distillation
[04:52] CTMs vs. variational score distillation
[05:20] Key takeaways for practitioners
[06:09] JVP rearrangement and flash attention
[06:52] FID metric evaluation
[07:32] Adaptive weighting benefits
[08:03] Future research directions
[08:37] Conclusion
Authors: Cheng Lu, Yang Song
Affiliations: OpenAI
Abstract: Consistency models (CMs) are a powerful class of diffusion-based generative models optimized for fast sampling. Most existing CMs are trained using discretized timesteps, which introduce additional hyperparameters and are prone to discretization errors. While continuous-time formulations can mitigate these issues, their success has been limited by training instability. To address this, we propose a simplified theoretical framework that unifies previous parameterizations of diffusion models and CMs, identifying the root causes of instability. Based on this analysis, we introduce key improvements in diffusion process parameterization, network architecture, and training objectives. These changes enable us to train continuous-time CMs at an unprecedented scale, reaching 1.5B parameters on ImageNet 512x512. Our proposed training algorithm, using only two sampling steps, achieves FID scores of 2.06 on CIFAR-10, 1.48 on ImageNet 64x64, and 1.88 on ImageNet 512x512, narrowing the gap in FID scores with the best existing diffusion models to within 10%.
From the publisher's feed