
Sign up to save your podcasts
Or


[00:00] Introduction to 3D Gaussian tracking for robotic manipulation
[00:26] Limitations of current video prediction methods
[01:11] Advantages of 3D Gaussian representation
[02:04] Graph Neural Networks for modeling object dynamics
[02:54] Control particle implementation and computation reduction
[03:42] Physics-based optimization for prediction stability
[04:25] Integration with real-world robotic systems
[05:12] Performance testing across different materials
[05:58] Advantages over traditional physics-based methods
[09:16] Implementation of object detection systems
[10:02] Data collection and synchronization challenges
[14:39] Long-term prediction capabilities and limitations
Authors: Mingtong Zhang, Kaifeng Zhang, Yunzhu Li
Affiliations: University of Illinois Urbana-Champaign, Columbia University
Abstract: Videos of robots interacting with objects encode rich information about the objects' dynamics. However, existing video prediction approaches typically do not explicitly account for the 3D information from videos, such as robot actions and objects' 3D states, limiting their use in real-world robotic applications. In this work, we introduce a framework to learn object dynamics directly from multi-view RGB videos by explicitly considering the robot's action trajectories and their effects on scene dynamics. We utilize the 3D Gaussian representation of 3D Gaussian Splatting (3DGS) to train a particle-based dynamics model using Graph Neural Networks. This model operates on sparse control particles downsampled from the densely tracked 3D Gaussian reconstructions. By learning the neural dynamics model on offline robot interaction data, our method can predict object motions under varying initial configurations and unseen robot actions. The 3D transformations of Gaussians can be interpolated from the motions of control particles, enabling the rendering of predicted future object states and achieving action-conditioned video prediction. The dynamics model can also be applied to model-based planning frameworks for object manipulation tasks. We conduct experiments on various kinds of deformable materials, including ropes, clothes, and stuffed animals, demonstrating our framework's ability to model complex shapes and dynamics. Our project page is available at this https URL.
Link: https://arxiv.org/abs/2410.18912
[00:00] Introduction
[00:20] Core limitations in robot manipulation: challenges with RL and IL
[01:08] SPIRE's hybrid approach: combining task planning with learning methods
[01:44] TAMP-gated learning: selective application of learned policies
[02:20] Training innovations: warm-starting RL and KL-divergence implementation
[02:59] Results: 35-50% performance gain, 6x more data efficient
[04:04] Multi-worker framework: improved sampling and distribution
[05:11] Future directions: expanding beyond rigid objects
[05:59] Curriculum learning: sequential training strategies
[07:11] Safety improvements: demonstrated through coffee task example
Authors: Zihan Zhou, Animesh Garg, Dieter Fox, Caelan Garrett, Ajay Mandlekar
Affiliations: NVIDIA, University of Toronto, Vector Institute, Georgia Institute of Technology
Abstract: Robot learning has proven to be a general and effective technique for programming manipulators. Imitation learning is able to teach robots solely from human demonstrations but is bottlenecked by the capabilities of the demonstrations. Reinforcement learning uses exploration to discover better behaviors; however, the space of possible improvements can be too large to start from scratch. And for both techniques, the learning difficulty increases proportional to the length of the manipulation task. Accounting for this, we propose SPIRE, a system that first uses Task and Motion Planning (TAMP) to decompose tasks into smaller learning subproblems and second combines imitation and reinforcement learning to maximize their strengths. We develop novel strategies to train learning agents when deployed in the context of a planning system. We evaluate SPIRE on a suite of long-horizon and contact-rich robot manipulation problems. We find that SPIRE outperforms prior approaches that integrate imitation learning, reinforcement learning, and planning by 35% to 50% in average task performance, is 6 times more data efficient in the number of human demonstrations needed to train proficient agents, and learns to complete tasks nearly twice as efficiently. View this https URL for more details.
Link: https://arxiv.org/abs/2410.18065
[00:00] VILA-U: A unified visual AI model
[00:29] Problem: Inefficiency of separate visual modules
[01:11] Vision tower: Novel quantization approach
[02:09] Training strategy: CLIP-based staged learning
[03:03] RVQ technique: Enhanced visual representation
[03:47] Multi-modal training: Text-image-video fusion
[04:35] Performance: Results and current limitations
[05:23] Impact: Contrastive loss effectiveness
[06:03] Generation: Optimal guidance settings
[06:37] Capabilities: Video, Q&A, and image reasoning
[07:14] Applications: Future use cases and scaling
[08:00] Architecture: LLaMA 2 7B integration
[08:48] Data: Quality vs quantity considerations
[09:35] Impact: Unified framework achievements
Authors: Yecheng Wu, Zhuoyang Zhang, Junyu Chen, Haotian Tang, Dacheng Li, Yunhao Fang, Ligeng Zhu, Enze Xie, Hongxu Yin, Li Yi, Song Han, Yao Lu
Affiliations: Tsinghua University, MIT, NVIDIA, UC Berkeley, UC San Diego
Abstract: VILA-U is a Unified foundation model that integrates Video, Image, Language understanding and generation. Traditional visual language models (VLMs) use separate modules for understanding and generating visual content, which can lead to misalignment and increased complexity. In contrast, VILA-U employs a single autoregressive next-token prediction framework for both tasks, eliminating the need for additional components like diffusion models. This approach not only simplifies the model but also achieves near state-of-the-art performance in visual language understanding and generation. The success of VILA-U is attributed to two main factors: the unified vision tower that aligns discrete visual tokens with textual inputs during pretraining, which enhances visual perception, and autoregressive image generation can achieve similar quality as diffusion models with high-quality dataset. This allows VILA-U to perform comparably to more complex models using a fully token-based autoregressive framework.
Link: https://arxiv.org/abs/2409.04429
[00:00] Intro
[00:31] Challenge: Limited multi-humanoid training data
[00:55] CooHOI's two-phase learning framework
[01:49] Object dynamics as implicit agent communication
[02:25] Bounding box strategy for long objects
[03:07] Results: Superior performance vs baselines
[03:42] Ablation study findings: Key system components
[04:18] Limitation: Basic hand manipulation only
[04:57] Impact: New approach to robot cooperation
Authors: Jiawei Gao, Ziqin Wang, Zeqi Xiao, Jingbo Wang, Tai Wang, Jinkun Cao, Xiaolin Hu, Si Liu, Jifeng Dai, Jiangmiao Pang
Affiliations: Shanghai AI Laboratory, Tsinghua University, Beihang University, Nanyang Technological University, Carnegie Mellon University
Abstract: Recent years have seen significant advancements in humanoid control, largely due to the availability of large-scale motion capture data and the application of reinforcement learning methodologies. However, many real-world tasks, such as moving large and heavy furniture, require multi-character collaboration. Given the scarcity of data on multi-character collaboration and the efficiency challenges associated with multi-agent learning, these tasks cannot be straightforwardly addressed using training paradigms designed for single-agent scenarios. In this paper, we introduce Cooperative Human-Object Interaction (CooHOI), a novel framework that addresses multi-character objects transporting through a two-phase learning paradigm: individual skill acquisition and subsequent transfer. Initially, a single agent learns to perform tasks using the Adversarial Motion Priors (AMP) framework. Following this, the agent learns to collaborate with others by considering the shared dynamics of the manipulated object during parallel training using Multi Agent Proximal Policy Optimization (MAPPO). When one agent interacts with the object, resulting in specific object dynamics changes, the other agents learn to respond appropriately, thereby achieving implicit communication and coordination between teammates. Unlike previous approaches that relied on tracking-based methods for multi-character HOI, CooHOI is inherently efficient, does not depend on motion capture data of multi-character interactions, and can be seamlessly extended to include more participants and a wide range of object types.
Link: https://arxiv.org/abs/2406.14558v2
[00:00] Introduction to SynFlowNet
[00:29] Problem: AI-generated molecules often can't be synthesized
[01:17] Solution: SynFlowNet - uses real chemical reactions
[02:03] GFlowNets: Enables diverse molecule generation
[02:47] Scalability: Morgan fingerprints handle 200K+ compounds
[03:14] Challenge: Solving backward trajectory issues
[04:14] Results: Better synthesis rates and molecular diversity
[05:30] Scale test: Successfully handled 221K molecules
[06:06] Application: Integration with fragment screening
[06:38] Wrap-up: SynFlowNet advances drug design
Authors: Miruna Cretu, Charles Harris, Ilia Igashov, Arne Schneuing, Marwin Segler, Bruno Correia, Julien Roy, Emmanuel Bengio, Pietro Liò
Affiliations: University of Cambridge, EPFL, Microsoft Research, Valence Labs
Abstract: Generative models see increasing use in computer-aided drug design. However, while performing well at capturing distributions of molecular motifs, they often produce synthetically inaccessible molecules. To address this, we introduce SynFlowNet, a GFlowNet model whose action space uses chemical reactions and buyable reactants to sequentially build new molecules. By incorporating forward synthesis as an explicit constraint of the generative mechanism, we aim at bridging the gap between in silico molecular generation and real world synthesis capabilities. We evaluate our approach using synthetic accessibility scores and an independent retrosynthesis tool to assess the synthesizability of our compounds, and motivate the choice of GFlowNets through considerable improvement in sample diversity compared to baselines. Additionally, we identify challenges with reaction encodings that can complicate traversal of the MDP in the backward direction. To address this, we introduce various strategies for learning the GFlowNet backward policy and thus demonstrate how additional constraints can be integrated into the GFlowNet MDP framework. This approach enables our model to successfully identify synthesis pathways for previously unseen molecules.
Link: https://arxiv.org/abs/2405.01155v2
[00:00] Intro to L3DG for 3D modeling
[00:32] Solving room-sized 3D scene complexity
[01:36] VQ-VAE compresses 3D Gaussian representation
[02:41] Generative sparse transpose convolution
[03:20] Latent diffusion for scene generation
[04:30] Visual improvements over baselines
[05:14] Scalability challenges for room-sized scenes
[06:13] Spherical harmonics for view dependence
[06:58] RGB and perceptual loss in training
[07:59] L1 and SSIM for 3D Gaussian optimization
[08:55] Training pipeline overview
[09:59] Densification in 3D Gaussian optimization
[10:49] Hyperparameter selection impact
[11:45] Future research directions
[12:42] Implementation optimization potential
[13:29] Comparison with GANs and diffusion methods
[14:27] Sparse grid representation trade-offs
[15:26] Evaluation datasets
[16:20] Chamfer distance for geometric analysis
[17:13] Applications
Authors: Barbara Roessle, Norman Müller, Lorenzo Porzi, Samuel Rota Bulò, Peter Kontschieder, Angela Dai, Matthias Nießner
Affiliations: Technical University of Munich, Meta Reality Labs Zurich
Abstract: We propose L3DG, the first approach for generative 3D modeling of 3D Gaussians through a latent 3D Gaussian diffusion formulation. This enables effective generative 3D modeling, scaling to generation of entire room-scale scenes which can be very efficiently rendered. To enable effective synthesis of 3D Gaussians, we propose a latent diffusion formulation, operating in a compressed latent space of 3D Gaussians. This compressed latent space is learned by a vector-quantized variational autoencoder (VQ-VAE), for which we employ a sparse convolutional architecture to efficiently operate on room-scale scenes. This way, the complexity of the costly generation process via diffusion is substantially reduced, allowing higher detail on object-level generation, as well as scalability to large scenes. By leveraging the 3D Gaussian representation, the generated scenes can be rendered from arbitrary viewpoints in real-time. We demonstrate that our approach significantly improves visual quality over prior work on unconditional object-level radiance field synthesis and showcase its applicability to room-scale scene generation.
Link: https://arxiv.org/abs/2410.13530
[00:00] Intro
[00:33] Combining transformers & diffusion models
[01:12] Key design: Scalable attention blocks (AdaLN)
[02:30] Efficient observation tokenization
[03:45] DiT Block policy architecture overview
[04:20] BiPlay dataset introduction
[04:53] Performance improvements over baselines
[05:30] Key findings from ablations
[06:08] Generalization to different robot types
[06:43] Simulation vs real-world performance
[07:13] Takeaways and future research directions
Authors: Sudeep Dasari, Oier Mees, Sebastian Zhao, Mohan Kumar Srirama, Sergey Levine
Affiliations: Carnegie Mellon University, University of California, Berkeley.
Abstract: In recent years roboticists have achieved remarkable progress in solving increasingly general tasks on dexterous robotic hardware by leveraging high capacity Transformer network architectures and generative diffusion models. Unfortunately, combining these two orthogonal improvements has proven surprisingly difficult, since there is no clear and well-understood process for making important design choices. In this paper, we identify, study and improve key architectural design decisions for high-capacity diffusion transformer policies. The resulting models can efficiently solve diverse tasks on multiple robot embodiments, without the excruciating pain of per-setup hyper-parameter tuning. By combining the results of our investigation with our improved model components, we are able to present a novel architecture, named \method, that significantly outperforms the state of the art in solving long-horizon (1500+ time-steps) dexterous tasks on a bi-manual ALOHA robot. In addition, we find that our policies show improved scaling performance when trained on 10 hours of highly multi-modal, language annotated ALOHA demonstration data. We hope this work will open the door for future robot learning techniques that leverage the efficiency of generative diffusion modeling with the scalability of large scale transformer architectures. Code, robot dataset, and videos are available at: this https URL
Link: https://arxiv.org/abs/2410.10088
[00:00] Introduction to EgoAllo system
[00:38] Challenges in egocentric motion estimation
[01:20] Importance of spatial/temporal invariance
[02:11] Comparison of conditioning parameterizations
[02:57] Integration of hand observations
[03:50] Global alignment phase
[04:28] Guidance losses in sampling
[05:03] Handling longer sequences
[05:35] Evaluation results
[06:30] System limitations and future work
[07:13] Implications for other egocentric tasks
[08:05] Advantages of diffusion models
[09:07] Use of synthetic datasets
[09:53] Promising research directions
[10:43] Impact on future motion capture systems
[11:41] Comparison to traditional methods
[12:31] Improved hand estimation accuracy
[13:25] SLAM data inaccuracies impact
[14:09] Levenberg-Marquardt optimizer usage
[15:14] Adapting to complex environments
Authors: Brent Yi, Vickie Ye, Maya Zheng, Lea Müller, Georgios Pavlakos, Yi Ma, Jitendra Malik, Angjoo Kanazawa
Affiliation: UC Berkeley, UT Austin
Abstract: We present EgoAllo, a system for human motion estimation from a head-mounted device. Using only egocentric SLAM poses and images, EgoAllo guides sampling from a conditional diffusion model to estimate 3D body pose, height, and hand parameters that capture the wearer's actions in the allocentric coordinate frame of the scene. To achieve this, our key insight is in representation: we propose spatial and temporal invariance criteria for improving model performance, from which we derive a head motion conditioning parameterization that improves estimation by up to 18%. We also show how the bodies estimated by our system can improve the hands: the resulting kinematic and temporal constraints result in over 40% lower hand estimation errors compared to noisy monocular estimates.
Project page: https://egoallo.github.io/
[00:00] Intro
[00:28] Limitation of existing unified models
[00:57] Janus's decoupled visual encoding solution
[01:18] Advantages of decoupling
[02:03] Janus architecture
[02:50] Three-stage training
[03:41] Ablation studies
[04:23] Extensions for Janus
[05:10] Performance gains
[05:47] Current limitations
[06:31] Impact of simplicity and extensibility
[07:10] Qualitative results
[08:18] Potential applications
[08:52] Key takeaways
Abstract: In this paper, we introduce Janus, an autoregressive framework that unifies multimodal understanding and generation. Prior research often relies on a single visual encoder for both tasks, such as Chameleon. However, due to the differing levels of information granularity required by multimodal understanding and generation, this approach can lead to suboptimal performance, particularly in multimodal understanding. To address this issue, we decouple visual encoding into separate pathways, while still leveraging a single, unified transformer architecture for processing. The decoupling not only alleviates the conflict between the visual encoder's roles in understanding and generation, but also enhances the framework's flexibility. For instance, both the multimodal understanding and generation components can independently select their most suitable encoding methods. Experiments show that Janus surpasses previous unified model and matches or exceeds the performance of task-specific models. The simplicity, high flexibility, and effectiveness of Janus make it a strong candidate for next-generation unified multimodal models.
Authors: Chengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, Chong Ruan, Ping Luo
Affiliations: DeepSeek-AI, The University of Hong Kong, Peking University
Link: https://arxiv.org/abs/2410.13848
[00:00] Introduction
[00:23] Computational cost of traditional diffusion models
[00:59] Reducing iterations in image generation
[01:06] Shortcut models
[01:39] Training process and self-consistency property
[02:22] Advantages over other methods
[03:05] Results on image generation benchmarks
[03:45] Application to robotic control
[04:15] Limitations and future work
[04:54] Best practices
Authors: Kevin Frans, Danijar Hafner, Sergey Levine, Pieter Abbeel
Affiliations: UC Berkeley
Abstract: Diffusion models and flow-matching models have enabled generating diverse and realistic images by learning to transfer noise to data. However, sampling from these models involves iterative denoising over many neural network passes, making generation slow and expensive. Previous approaches for speeding up sampling require complex training regimes, such as multiple training phases, multiple networks, or fragile scheduling. We introduce shortcut models, a family of generative models that use a single network and training phase to produce high-quality samples in a single or multiple sampling steps. Shortcut models condition the network not only on the current noise level but also on the desired step size, allowing the model to skip ahead in the generation process. Across a wide range of sampling step budgets, shortcut models consistently produce higher quality samples than previous approaches, such as consistency models and reflow. Compared to distillation, shortcut models reduce complexity to a single network and training phase and additionally allow varying step budgets at inference time.
Link: https://arxiv.org/abs/2410.12557
From the publisher's feed