
Sign up to save your podcasts
Or


Google is preparing to release their latest AI software Gemini, which aims to compete with OpenAI's GPT-4 model, and can be used for everything from chatbots to generating original text, music lyrics, and news stories. "Generative Image Dynamics" is a paper from Google Research that focuses on creating a model for scene dynamics in images, which can be used to turn still images into seamlessly looping dynamic videos or allow users to realistically interact with objects in real pictures. "Agents: An Open-source Framework for Autonomous Language Agents" is an open-source library that makes it easier for non-specialists to build and deploy state-of-the-art autonomous language agents that can interact with humans, other agents, and environments using natural language interfaces. "ExpertQA: Expert-Curated Questions and Attributed Answers" is a paper that presents an evaluation study that analyzes factuality and attribution in responses from language models in domain-specific scenarios, resulting in a high-quality long-form QA dataset that can be used for various applications like building better language models or training AI systems for specific domains.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:47 Google nears release of AI software Gemini
03:25 Funky AI-generated spiraling medieval village captivates social media
05:13 Jason Wei Tweets on Pair Programming
06:23 Fake sponsor
08:19 Agents: An Open-source Framework for Autonomous Language Agents
09:50 Generative Image Dynamics
11:07 ExpertQA: Expert-Curated Questions and Attributed Answers
13:12 Outro
Stable Audio, a new text-to-audio generative AI platform, uses a diffusion model trained with audio to create background music for podcasts or videos. Amazon has launched a new generative AI tool to help sellers write better product descriptions, which dramatically improves the listing creation and management experience for sellers. OpenAI is opening an office in Dublin, Ireland, to collaborate with the government and industry to advance AI development and deployment. Three papers were discussed, including the use of Large Language Models (LLMs) to optimize code, MagiCapture's personalization method for generating high-resolution portrait images, and Statistical Rejection Sampling Optimization (RSO) to improve language models' alignment with human preferences.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:17 Stability AI releases AI audio platform
03:00 Amazon launches generative AI to help sellers write product descriptions
04:41 OpenAI Opening an Office in Dublin
06:12 Fake sponsor
08:16 Large Language Models for Compiler Optimization
09:25 MagiCapture: High-Resolution Multi-Concept Portrait Customization
11:20 Statistical Rejection Sampling Improves Preference Optimization
13:13 Outro
Apple's use of AI in its new devices, including a new chip that includes improved data crunching capabilities and a four-core "Neural Engine" that can process machine learning tasks up to twice as quickly. Elon Musk's proposal for a federal department of AI after his Capitol Hill summit, citing the potential harm of unchecked AI development. Cutting-edge research on efficient memory management, mesa-optimization algorithms, and extremely parameter-efficient MoE. The proposed solutions to challenges in serving large language models efficiently, including PagedAttention and vLLM.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:41 AI quietly reshapes Apple iPhones, Watches
03:06 Elon Musk calls for federal department of AI after Capitol Hill summit
04:34 Jason Wei tweets
06:15 Fake sponsor
08:06 Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction Tuning
09:33 Uncovering mesa-optimization algorithms in Transformers
11:06 Efficient Memory Management for Large Language Model Serving with PagedAttention
12:53 Outro
NVIDIA is teaming up with leaders to discuss AI standards, Coca-Cola has released a new zero sugar flavor created with AI, Microsoft Research has introduced a new 1.3 billion parameter model named phi-1.5, and Google DeepMind and Google Research have introduced MADLAD-400, a manually audited, general domain 3T token monolingual dataset based on CommonCrawl, spanning 419 languages. Tune in to hear more about these exciting developments in the world of AI.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:57 NVIDIA Lends Support to Washington’s Efforts to Ensure AI Safety
03:30 Coca-Cola Drops a Zero Sugar Flavor Created by AI
04:39 Nvidia Shows Off Grace Hopper in MLPerf Inference
06:31 Fake sponsor
08:37 Textbooks Are All You Need II: phi-1.5 technical report
10:18 MADLAD-400: A Multilingual And Document-Level Large Audited Dataset
11:56 Robot Parkour Learning
13:35 Outro
G20's reaffirmation of responsible AI use, Meta's plans for a new chatbot model, Google DeepMind's explanation for grokking in neural networks, and a system for automatically generating high-quality audiobooks from online e-books.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:36 G20 nations reaffirm responsible use and development of AI technology
03:13 Meta sets GPT-4 as the bar for its next AI model, says a new report
04:48 Deep Neural Nets: 33 years ago and 33 years from now
06:29 Fake sponsor
08:25 Explaining grokking through circuit efficiency
10:02 Large-Scale Automatic Audiobook Creation
11:39 Subwords as Skills: Tokenization for Sparse-Reward Reinforcement Learning
13:13 Outro
From the Pentagon's plans for a vast AI fleet to counter the China threat, to OpenAI's confirmation that AI writing detectors just don't cut it, we cover it all. We also explore the emergent abilities in large language models and their potential to revolutionize optimization in various fields, as well as ImageBind-LLM, a new method for tuning large language models with multi-modality instructions.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:46 Pentagon Plans Vast AI Fleet to Counter China Threat
03:01 OpenAI confirms that AI writing detectors don’t work
04:46 Why is the ocean salty?
05:48 Fake sponsor
07:26 Are Emergent Abilities in Large Language Models just In-Context Learning?
08:43 Large Language Models as Optimizers
10:17 ImageBind-LLM: Multi-modality Instruction Tuning
12:28 Outro
The episode covers cutting-edge AI research on vision and language models, including a new pretraining methodology for open-vocabulary object detection and a physically grounded VLM for robotic manipulation tasks. The show also features two interesting papers on DSPy, a framework for working with language models and retrieval models, and Verba, an open-source initiative for retrieval-augmented generation applications. The crew discusses the TIME100 Most Influential People in AI, highlighting the significance of generative AI and the ethical questions surrounding its development.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:38 How We Chose the TIME100 Most Influential People in AI
03:16 DSPy: Programming—not prompting—Foundation Models
04:28 Verba Retrieval Augmented Generation from Weaviate
05:26 Fake sponsor
07:28 Contrastive Feature Masking Open-Vocabulary Vision Transformer
09:24 Physically Grounded Vision-Language Models for Robotic Manipulation
11:27 Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning
13:40 Outro
OpenAI announces their first developer conference, Zoom debuts an AI assistant, and we explore the discovery that certain RNNs might be implementing attention under the hood. We also discuss Sequential Dexterity, a system that chains multiple dexterous policies for achieving long-horizon task goals.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:37 OpenAI to hold its first developer conference on November 6 in San Francisco
03:03 Zoom Debuts AI assistant
04:39 Jason Wei Tweets about alternative altmetrics
06:28 Fake sponsor
08:44 Gated recurrent neural networks discover attention
10:07 One Wide Feedforward is All You Need
11:48 Sequential Dexterity: Chaining Dexterous Policies for Long-Horizon Manipulation
13:26 Outro
Amazon One for convenient payment and verification, the concerning trend of AI-generated sex workers flooding social media platforms, OpenAI's legal battle over ChatGPT's alleged use of pirated books, and a promising solution to the challenge of catastrophic forgetting in Continual Learning with LGCL.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:33 How generative AI helped train Amazon One to recognize your palm
02:55 Ads for AI sex workers are flooding Instagram and TikTok
04:07 OpenAI disputes authors’ claims that every ChatGPT response is a derivative work
05:39 Fake sponsor
07:33 Introducing Language Guidance in Prompt-based Continual Learning
09:20 Baseline Defenses for Adversarial Attacks Against Aligned Language Models
10:46 BatchPrompt: Accomplish more with less
12:31 Outro
Twitter's updated privacy policy and how they plan to use public data to train their AI models. We also dive into OpenAI's Guide to Teaching with AI and explore the potential benefits and limitations of using AI in education. Additionally, we highlight some cutting-edge research papers on large language and speech models, a unified speech tokenizer, and a multimodal wine dataset. And, for a bit of fun, we have an entertaining ad for SplashTech's SuperSoak water gun.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:41 Twitter’s privacy policy confirms it will use public data to train AI models
02:56 OpenAI's Guide to Teaching with AI
04:33 Introducing Refact Code LLM: 1.6B State-of-the-Art LLM for Code that Reaches 32% HumanEval
05:48 Fake sponsor
08:20 LLaSM: Large Language and Speech Model
09:50 SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
11:52 Learning to Taste: A Multimodal Wine Dataset
13:36 Outro
From the publisher's feed