AI Illuminated

AI Illuminated

By The AI IlluminatorsEducationCourses
Download on the App Store

AI Illuminated episodes

  • Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning

    [00:00] Intro: Research on multi-task learning in LLMs

    [00:38] Balancing safety and performance in multilingual settings

    [01:17] Model merging techniques explored

    [02:16] Model merging outperforms data mixing

    [02:49] Merging monolingual models improves multilingual capabilities

    [03:28] Key ablation studies

    [04:14] Safety and performance evaluation metrics

    [04:52] Effectiveness variations across languages

    [05:26] Safety model weight impact in linear merging

    [06:00] Insights on merging and preference training

    [06:25] Comparison to existing research

    [06:56] Implications for LLM development

    [07:29] Limitations and future research


    Authors: Aakanksha, Arash Ahmadian, Seraphina Goldfarb-Tarrant, Beyza Ermis, Marzieh Fadaee, Sara Hooker

    Affiliations: Cohere For AI, Cohere

    Abstract: Large Language Models (LLMs) have been adopted and deployed worldwide for a broad variety of applications. However, ensuring their safe use remains a significant challenge. Preference training and safety measures often overfit to harms prevalent in Western-centric datasets, and safety protocols frequently fail to extend to multilingual settings. In this work, we explore model merging in a diverse multi-task setting, combining safety and general-purpose tasks within a multilingual context. Each language introduces unique and varied learning challenges across tasks. We find that objective-based merging is more effective than mixing data, with improvements of up to 8% and 10% in general performance and safety respectively. We also find that language-based merging is highly effective -- by merging monolingually fine-tuned models, we achieve a 4% increase in general performance and 7% reduction in harm across all languages on top of the data mixtures method using the same available data. Overall, our comprehensive study of merging approaches provides a useful framework for building strong and safe multilingual models.

    Link: https://arxiv.org/abs/2410.10801

    8 min
  • CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos

    [00:00] Intro: CoTracker 3 paper

    [00:17] Challenge: Point tracker training

    [00:51] Existing methods' limitations

    [01:05] CoTracker 3 innovations

    [01:41] Novel semi-supervised training

    [02:24] Online vs offline versions

    [03:02] Performance improvements

    [03:47] Ablation study insights

    [04:29] Limitations and future work


    Authors: Nikita Karaev, Iurii Makarov, Jianyuan Wang, Natalia Neverova, Andrea Vedaldi, Christian Rupprecht

    Affiliations: Meta AI; Visual Geometry Group, University of Oxford

    Abstract: Most state-of-the-art point trackers are trained on synthetic data due to the difficulty of annotating real videos for this task. However, this can result in suboptimal performance due to the statistical gap between synthetic and real videos. In order to understand these issues better, we introduce CoTracker3, comprising a new tracking model and a new semi-supervised training recipe. This allows real videos without annotations to be used during training by generating pseudo-labels using off-the-shelf teachers. The new model eliminates or simplifies components from previous trackers, resulting in a simpler and often smaller architecture. This training scheme is much simpler than prior work and achieves better results using 1,000 times less data. We further study the scaling behaviour to understand the impact of using more real unsupervised data in point tracking. The model is available in online and offline variants and reliably tracks visible and occluded points.

    Link: https://arxiv.org/abs/2410.11831

    6 min
  • nGPT: Normalized Transformer with Representation Learning on the Hypersphere

    [00:00] Introduction

    [00:30] Consistent unit norm normalization in NGPT

    [01:08] Mathematical mechanism behind faster convergence

    [01:52] Elimination of weight decay in NGPT

    [02:21] Role of learnable eigen learning rates in optimization

    [03:04] Discussion on training speedup vs. per-step computation time

    [03:46] Condition number differences between GPT and NGPT

    [04:18] Ablation studies on scaling factors

    [04:53] NGPT's relationship to Riemannian optimization

    [05:27] Future research

    [06:02] Takeaways for practitioners


    Authors: Ilya Loshchilov, Cheng-Ping Hsieh, Simeng Sun, Boris Ginsburg

    Affiliations: NVIDIA

    Abstract: We propose a novel neural network architecture, the normalized Transformer (nGPT) with representation learning on the hypersphere. In nGPT, all vectors forming the embeddings, MLP, attention matrices and hidden states are unit norm normalized. The input stream of tokens travels on the surface of a hypersphere, with each layer contributing a displacement towards the target output predictions. These displacements are defined by the MLP and attention blocks, whose vector components also reside on the same hypersphere. Experiments show that nGPT learns much faster, reducing the number of training steps required to achieve the same accuracy by a factor of 4 to 20, depending on the sequence length.




    7 min
  • RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

    [00:00] Intro

    [00:20] Challenges in bimanual manipulation models
    [01:00] Robotics Diffusion Transformer (RDT) approach
    [01:34] RDT architecture design
    [02:22] Data scarcity and unified action space
    [03:07] Multi-task bimanual dataset
    [03:49] RDT's experimental results
    [04:30] Benefits of large-scale pre-training
    [05:12] Diffusion models in robotics
    [06:02] RDT's architectural modifications
    [06:57] Unified action space benefits
    [07:46] GPT-4 Turbo for data augmentation
    [08:24] Real robot experiments
    [09:14] Future research directions

    Authors: Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan, Huayu Chen, Zhengyi Wang, Ke Xu, Hang Su, Jun Zhu

    Affiliations: Tsinghua University

    Abstract: Bimanual manipulation is essential in robotics, yet developing foundation models is extremely challenging due to the inherent complexity of coordinating two robot arms (leading to multi-modal action distributions) and the scarcity of training data. In this paper, we present the Robotics Diffusion Transformer (RDT), a pioneering diffusion foundation model for bimanual manipulation. RDT builds on diffusion models to effectively represent multi-modality, with innovative designs of a scalable Transformer to deal with the heterogeneity of multi-modal inputs and to capture the nonlinearity and high frequency of robotic data. To address data scarcity, we further introduce a Physically Interpretable Unified Action Space, which can unify the action representations of various robots while preserving the physical meanings of original actions, facilitating learning transferrable physical knowledge. With these designs, we managed to pre-train RDT on the largest collection of multi-robot datasets to date and scaled it up to 1.2B parameters, which is the largest diffusion-based foundation model for robotic manipulation. We finally fine-tuned RDT on a self-created multi-task bimanual dataset with over 6K+ episodes to refine its manipulation capabilities. Experiments on real robots demonstrate that RDT significantly outperforms existing methods. It exhibits zero-shot generalization to unseen objects and scenes, understands and follows language instructions, learns new skills with just 1~5 demonstrations, and effectively handles complex, dexterous tasks. We refer to this https URL for the code and videos.


    Link: https://rdt-robotics.github.io/rdt-robotics/

    10 min
  • Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models

    [00:00] Continuous-time consistency models (CTMs)

    [00:20] Limitations of existing CTMs

    [00:51] TrigFlow: New CTM formulation

    [01:42] CTM training instability

    [02:21] Training objective modifications

    [02:55] Scaling CTMs to 1.5 billion parameters

    [03:37] Comparison with state-of-the-art models

    [04:14] Consistency training vs. distillation

    [04:52] CTMs vs. variational score distillation

    [05:20] Key takeaways for practitioners

    [06:09] JVP rearrangement and flash attention

    [06:52] FID metric evaluation

    [07:32] Adaptive weighting benefits

    [08:03] Future research directions

    [08:37] Conclusion


    Authors: Cheng Lu, Yang Song

    Affiliations: OpenAI

    Abstract: Consistency models (CMs) are a powerful class of diffusion-based generative models optimized for fast sampling. Most existing CMs are trained using discretized timesteps, which introduce additional hyperparameters and are prone to discretization errors. While continuous-time formulations can mitigate these issues, their success has been limited by training instability. To address this, we propose a simplified theoretical framework that unifies previous parameterizations of diffusion models and CMs, identifying the root causes of instability. Based on this analysis, we introduce key improvements in diffusion process parameterization, network architecture, and training objectives. These changes enable us to train continuous-time CMs at an unprecedented scale, reaching 1.5B parameters on ImageNet 512x512. Our proposed training algorithm, using only two sampling steps, achieves FID scores of 2.06 on CIFAR-10, 1.48 on ImageNet 64x64, and 1.88 on ImageNet 512x512, narrowing the gap in FID scores with the best existing diffusion models to within 10%.

    10 min

About AI Illuminated

From the publisher's feed

A new way to keep up with AI research. Delivered to your ears. Illuminated by AI.