
Sign up to save your podcasts
Or


Nvidia's Q1 revenue up 262% to $26.0B, beating estimates.
OpenAI's News Corp deal licenses content from WSJ, New York Post and more.
PyramidInfer compresses KV cache to save memory during inference for Large Language Models.
Your Transformer is Secretly Linear challenges our existing understanding of transformer architectures.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:55 Nvidia's Q1 revenue up 262% to $26.0B, beating estimates
03:23 OpenAI’s News Corp deal licenses content from WSJ, New York Post, and more
04:57 Systematically Improving Your RAG
06:18 Fake sponsor
08:17 PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
09:49 Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
11:48 Your Transformer is Secretly Linear
13:26 Outro
Google's redesign of its search engine using AI to enhance the search experience.
Microsoft's introduction of Copilot+, the first AI PC, with innovative features like Recall and Cocreator.
Imp, a highly capable large multimodal model for mobile devices that can process and understand multiple types of data simultaneously.
Octo, an open-source generalist robot policy that can be finetuned to new observation and action spaces, potentially transforming robotic learning.
Contact: [email protected]
Timestamps:
00:34 Introduction
02:14 Google is redesigning its search engine — and it’s AI all the way down
03:35 Microsoft Unveils First AI PC
05:18 Greg Brockman and Sam Altman Statement on OpenAI's Aligmnent Schism
06:53 Fake sponsor
08:55 Imp: Highly Capable Large Multimodal Models for Mobile Devices
10:37 Octo: An Open-Source Generalist Robot Policy
12:11 Towards Modular LLMs by Building and Reusing a Library of LoRAs
14:02 Outro
OpenAI's ChatGPT introduces new enhancements for data analysis, making it easier for beginners to perform in-depth analyses and saves experts time on routine data-cleaning tasks.
Sony Music warns AI companies against unauthorized use of their copyrighted material for the "training, development or commercialization of AI systems", highlighting concerns around the use of AI-generated voice clones.
Chameleon, a family of models that can understand and generate both images and text in any sequence, uses an early-fusion approach, resulting in better performance across a wide range of tasks.
MoRA proposes a new method for fine-tuning large language models, achieving high-rank updating while maintaining the same number of trainable parameters, which could have practical implications for improving the efficiency of large language models.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:58 Improvements to data analysis in ChatGPT
03:31 Sony Music warns AI companies against ‘unauthorized use’ of its content
05:28 Statement from Scarlett Johansson on the OpenAI situation
06:52 Fake sponsor
09:02 Chameleon: Mixed-Modal Early-Fusion Foundation Models
10:27 Layer-Condensed KV Cache for Efficient Inference of Large Language Models
12:00 MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning
13:47 Outro
OpenAI team imploding due to a loss of faith in leadership and prioritization of safety over commercialization.
Reddit and OpenAI partnership to bring Reddit content to ChatGPT and introduce new AI-powered features to users.
The extensive process OpenAI went through to select the five distinct voices for ChatGPT's Voice Mode.
Papers discussing Layer-Condensed KV Cache for efficient inference of large language models, observational scaling laws and the predictability of language model performance, and Chameleon's mixed-modal early-fusion foundation models.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:27 Why the OpenAI team in charge of safeguarding humanity imploded
03:03 Reddit and OpenAI Build Partnership
04:33 How the voices for ChatGPT were chosen
05:54 Fake sponsor
07:22 Layer-Condensed KV Cache for Efficient Inference of Large Language Models
08:48 Observational Scaling Laws and the Predictability of Language Model Performance
10:35 Chameleon: Mixed-Modal Early-Fusion Foundation Models
12:20 Outro
Google I/O 2024 announcements, including new AI tools like Firebase Genkit, LearnLM, and Veo, as well as Gemini, an AI replacement for Google Assistant.
The introduction of the MS MARCO Web Search dataset, which provides a retrieval benchmark with three web retrieval challenge tasks and millions of real-clicked query-document pairs for training and evaluating retrieval models.
The "What matters when building vision-language models?" paper, which identifies critical decisions in the design of vision-language models and presents Idefics2, an efficient foundational VLM of 8 billion parameters that achieves state-of-the-art performance within its size category.
The "RLHF Workflow: From Reward Modeling to Online RLHF" paper, which presents a workflow for Online Iterative Reinforcement Learning from Human Feedback (RLHF) in an online setting and achieves impressive performance on LLM chatbot benchmarks and academic benchmarks.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:34 Google I/O 2024: Here’s everything Google just announced
03:26 Ilya Sutskever leaves OpenAI
04:57 GPT-4o’s Memory Breakthrough!
06:00 Fake sponsor
07:49 MS MARCO Web Search: a Large-scale Information-rich Web Dataset with Millions of Real Click Labels
09:33 What matters when building vision-language models?
10:54 RLHF Workflow: From Reward Modeling to Online RLHF
13:00 Outro
OpenAI's new model, GPT-4o, can reason across audio, vision, and text in real-time, with safety measures built-in by design.
Apple and Google collaborate to deliver support for unwanted tracking alerts in iOS and Android, an industry first involving community and industry input.
LoRA Land, a web application that hosts 25 LoRA fine-tuned Mistral-7B LLMs on a single NVIDIA A100 GPU with 80GB memory, highlights the quality and cost-effectiveness of employing multiple specialized LLMs over a single, general-purpose LLM.
WildChat, a public dataset showcasing how chatbots like GPT-4 and ChatGPT are used by a population of users in practice, offers the most diverse user prompts, contains the largest number of languages, and presents the richest variety of potentially toxic use-cases for researchers to study.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:35 OpenAI Announces GPT-4 Omni
03:00 Apple and Google deliver support for unwanted tracking alerts in iOS and Android
05:03 Sam Altman on GPT-4 Omni
06:20 Fake sponsor
08:43 LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report
10:17 Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
12:04 WildChat: 1M ChatGPT Interaction Logs in the Wild
14:04 Outro
OpenAI plans to challenge Google search with a new search feature for ChatGPT, which could have a significant impact on the AI industry.
SoundHound AI and Perplexity have partnered to improve the accuracy and complexity of voice assistants across cars, apps, and phone assistants.
"Fishing for Magikarp" addresses an issue with "glitch tokens" in large language models, while "CuMo" proposes a new approach to improving multimodal LLMs.
"Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?" highlights the risk in introducing new factual knowledge through fine-tuning and the importance of pre-training for large language models.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:38 OpenAI plans to announce Google search competitor today
03:02 SoundHound AI and Perplexity Partner to Bring Online LLMs to Next Gen Voice Assistants Across Cars and IoT Devices
04:48 Homoiconic Python
06:06 Fake sponsor
07:49 Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
09:00 CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
10:34 Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
12:13 Outro
DeepMind's AlphaFold 3, the newest and most powerful version of their AI model that can predict the structure of proteins and other molecules with incredible accuracy, has been released for free for non-commercial use.
OpenAI has introduced the Model Spec, a document that specifies how they want their AI models to behave in their API and ChatGPT, to deepen the public conversation about how AI models should behave.
Microsoft Research's paper explores how players can interact with large language models (LLMs) to create emergent behaviors in game narratives, which could have big implications for game development and player engagement.
The University of California, Berkeley's paper proposes the Learnable Latent Codes as Bridges (LCB) method, which uses a learnable latent code as a bridge between LLMs and low-level policies, allowing for more flexible communication of goals in the task plan without being entirely constrained by language limitations.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:18 DeepMind Announces AlphaFold 3
02:37 OpenAI Introduces the Model Spec
04:22 ChatBotArena: The peoples’ LLM evaluation, the future of evaluation, the incentives of evaluation, and gpt2chatbot
05:53 Fake sponsor
08:11 Player-Driven Emergence in LLM-Driven Game Narrative
09:25 vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
11:15 From LLMs to Actions: Latent Codes as Bridges in Hierarchical Robot Control
12:55 Outro
Apple's new M4 chip for the iPad Pro promises improved performance and AI capabilities.
OpenAI is developing a search feature for ChatGPT that could rival Google and Perplexity.
IBM's Granite Language Models are a promising tool for code generative tasks, with improved performance and trustworthy data usage.
xLSTM and vAttention are two new approaches to optimizing the performance of large language models, with potential applications in natural language processing, speech recognition, and video analysis.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:42 Apple introduces M4 chip
03:34 IBM Granite Language Models
04:40 OpenAI Is Readying a Search Product to Rival Google, Perplexity
05:47 Fake sponsor
07:36 xLSTM: Extended Long Short-Term Memory
09:17 vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
10:46 NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
12:45 Outro
Stack Overflow and OpenAI partner to provide developers with accurate and vetted data for AI development.
Elon Musk plans to use AI to distill and present news on X, combining breaking news and social media reactions.
HuggingFace's Robotics Library, LeRobot, provides state-of-the-art machine learning models, datasets, and tools for real-world robotics.
Research papers on AI language retrieval explore improving multilingual information retrieval and the effects of downsizing large language models.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:33 Stack Overflow and OpenAI Partner to Strengthen the World’s Most Popular Large Language Models
03:21 Elon Musk's AI News Plans for X
05:15 LeRobot: HuggingFace's Robotics Library
06:27 Fake sponsor
08:22 Distillation for Multilingual Information Retrieval
10:03 The Cost of Down-Scaling Language Models: Fact Recall Deteriorates before In-Context Learning
11:35 In-Context Learning with Long-Context Models: An In-Depth Exploration
13:26 Outro
From the publisher's feed