
Sign up to save your podcasts
Or


AI expert Andrej Karpathy is starting an AI+Education company called Eureka Labs, which aims to create an ideal experience for learning something new by leveraging recent progress in generative AI.
YouTube Music is rolling out new AI-powered features, including 'Sound Search' that allows users to search YouTube's catalog of over 100 million songs by singing, humming, or playing a tune, and 'AI-generated conversational radio' that creates a tailored playlist based on natural language prompts.
"xLSTMTime: Long-term Time Series Forecasting with xLSTM" explores the use of extended LSTM (xLSTM) for improving long-term time series forecasting, demonstrating superior forecasting capabilities compared to other state-of-the-art models.
"Still-Moving: Customized Video Generation without Customized Video Data" introduces a framework called Still-Moving that seamlessly integrates the spatial prior of a customized text-to-image (T2I) model with the motion prior of a T2V model, achieving impressive results on personalized, stylized, and conditional video generation tasks.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:57 Andrej Karpathy starting an AI Education Company
03:27 YouTube Music sound search rolling out, AI ‘conversational radio’ in testing
05:12 GB200 Hardware Architecture & Component Supply Chain & BOM
06:50 Fake sponsor
08:34 xLSTMTime : Long-term Time Series Forecasting With xLSTM
10:04 Still-Moving: Customized Video Generation without Customized Video Data
11:47 NeedleBench: Can LLMs Do Retrieval and Reasoning in 1 Million Context Window?
13:32 Outro
OpenAI's mysterious code name 'Strawberry' and its potential to revolutionize AI capabilities.
The appearance of new models in the LMSYS Chatbot Arena, potentially hinting at a new mini-GPT release from OpenAI.
The introduction of EM-LLM, a new model that integrates human episodic memory and event cognition into LLMs, improving their ability to process extensive contexts.
The creation of SPIQA, a large-scale question answering dataset specifically designed to interpret complex figures and tables within the context of scientific research articles.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:44 Exclusive: OpenAI working on new reasoning technology under code name ‘Strawberry’
03:25 Mysterious Models in LMSys Arena
04:51 Run Cuda on AMD GPUs
06:14 Fake sponsor
07:52 Human-like Episodic Memory for Infinite Context LLMs
09:56 Toto: Time Series Optimized Transformer for Observability
11:28 SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers
13:21 Outro
Google DeepMind's latest research on robot navigation using the Gemini 1.5 Pro.
OpenAI's new five-tier system to track progress towards artificial general intelligence.
The development of AI-powered memory aids and the ethical implications of relying on technology for memory recall.
Microsoft's innovative encoding framework, SheetCompressor, for large language models to understand and reason about spreadsheets.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:43 Gemini 1.5 Pro’s long context window help robots navigate the world?
03:00 OpenAI Scale Ranks Progress Toward ‘Human-Level’ Problem Solving
04:47 Inside the AI memory machine
05:56 MambaVision: A Hybrid Mamba-Transformer Vision Backbone
07:01 Fake sponsor
09:09 Still-Moving: Customized Video Generation without Customized Video Data
10:41 SpreadsheetLLM: Encoding Spreadsheets for Large Language Models
12:51 Outro
Microsoft and Apple drop OpenAI Board plans due to increased regulatory scrutiny in the AI sector.
Research papers on enhancing mathematical reasoning capabilities of large language models and improving mathematical problem-solving capabilities in visual contexts using Multi-modal Large Language Models (MLLMs).
FlashAttention-3, an algorithm that speeds up attention mechanism in large language models by up to 2 times faster than previous versions, while maintaining accuracy with lower precision numbers.
Adaptive In-Context Learning, a technique that simplifies the overall machine learning pipeline, making it more accessible for more organizations.
Contact: [email protected]
Timestamps:
00:34 Introduction
02:01 Microsoft, Apple Drop OpenAI Board Plans as Scrutiny Grows
03:41 Reproducing GPT-2 in C and CUDA
04:48 Adaptive In-Context Learning
06:25 FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
07:53 Fake sponsor
10:06 Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
11:35 MAVIS: Mathematical Visual Instruction Tuning
13:38 Outro
a16z's Oxygen initiative to provide AI startups with access to GPUs at below-market rates in exchange for equity, potentially reshaping the AI VC landscape.
OpenAI's decision to block users in China from accessing its tools and services, potentially accelerating the development of homegrown models by Chinese AI companies.
AMD's acquisition of Silo AI for $665M to enhance its AI chip capabilities and compete against industry leader Nvidia.
Three exciting AI papers on the limitations of vision language models, a new approach for video instruction tuning, and a novel framework for multi-agent collaboration called the Internet of Agents.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:35 a16z is trying to keep AI alive with Oxygen initiative
03:14 Chinese developers scramble as OpenAI blocks access in China
04:52 AMD to acquire Finnish startup Silo AI for $665M to step up in AI race
06:18 Fake sponsor
08:00 Vision language models are blind
09:30 Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
11:17 Internet of Agents: Weaving a Web of Heterogeneous Agents for Collaborative Intelligence
13:37 Outro
The MNLU-Pro dataset is a more robust and challenging massive multi-task language understanding dataset that's tailored to more rigorously benchmark large language models' capabilities.
The Composable Interventions framework allows researchers to study the effects of using multiple interventions on a language model, and the order in which interventions are applied can have a significant impact on their effectiveness.
The MJ-Bench benchmark evaluates the effectiveness of different types of multimodal judges in providing feedback for text-to-image generation models, and the experiments reveal that close-source VLMs generally provide better feedback.
The Associative Recurrent Memory Transformer (ARMT) is an approach that combines transformer self-attention for local context with segment-level recurrence for storage of task-specific information distributed over a long context, and it sets a new performance record in the recent BABILong multi-task long-context benchmark.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:32 MNLU-Pro Release on HuggingFace Datasets
03:48 Extrinsic Hallucinations in LLMs
04:53 RouteLLM
06:13 Fake sponsor
08:14 Composable Interventions for Language Models
09:45 MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
11:31 Associative Recurrent Memory Transformer
13:30 Outro
OpenAI's internal messaging systems were breached last year, and sensitive details about the company's technology were stolen, raising questions about OpenAI's security protocols and how they handle these types of incidents.
Despite the popularity of generative AI chatbots like ChatGPT, Google Search's market dominance is actually growing, which may not bode well with antitrust regulators.
"Learning to (Learn at Test Time): RNNs with Expressive Hidden States" proposes a new class of sequence modeling layers called Test-Time Training (TTT) layers, which have the potential to be a powerful tool for sequence modeling.
"Finding Visual Task Vectors" introduces a technique called visual prompting, which has potential applications in areas such as robotics and autonomous systems, and provides an alternative to traditional supervised learning.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:38 A Hacker Stole OpenAI Secrets
03:05 ChatGPT might rule the AI chatbots — but it can't beat Google Search
04:52 MobileLLM
06:17 Fake sponsor
08:20 Learning to (Learn at Test Time): RNNs with Expressive Hidden States
10:17 Many-Shot In-Context Learning
11:53 Finding Visual Task Vectors
13:38 Outro
Moshi, the first real-time AI voice assistant with 70 different emotions and speaking styles, has been unveiled by French startup Kyutai.
ElevenLabs' Reader App now features "Iconic Voices" which uses AI-generated voices of late Hollywood stars to read text content within the app.
Google DeepMind's paper "On scalable oversight with weak LLMs judging strong LLMs" explores scalable oversight protocols using large language models (LLMs) to enable humans to supervise superhuman AI.
"Learning to (Learn at Test Time): RNNs with Expressive Hidden States" proposes a new approach to sequence modeling using Test-Time Training (TTT) layers, which make the hidden state a machine learning model itself.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:32 Unveiling of Moshi: the first voice-enabled AI openly accessible to all
02:38 ElevenLabs Ionic Voices
04:12 Your guide to AI: July 2024
05:25 Fake sponsor
07:17 On scalable oversight with weak LLMs judging strong LLMs
08:54 Reasoning in Large Language Models: A Geometric Perspective
10:17 Learning to (Learn at Test Time): RNNs with Expressive Hidden States
12:12 Outro
HuggingFace has upgraded the Open LLM Leaderboard to v2, adding new benchmarks and improving the evaluation suite for easier reproducibility.
Gemma 2, a new addition to the Gemma family of lightweight open models, delivers the best performance for its size and offers competitive alternatives to models that are 2-3× bigger.
SeaKR is a new model that re-ranks retrieved knowledge based on the LLM's self-aware uncertainty, outperforming existing adaptive RAG methods in generating text with relevant and accurate information.
Step-DPO is a new method that enhances the robustness and factuality of LLMs by learning from human feedback, achieving impressive results in long-chain mathematical reasoning.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:21 HuggingFace Updates Open LLM Leaderboard
03:19 Gemma 2: Improving Open Language Models at a Practical Size
04:16 From bare metal to a 70B model: infrastructure set-up and scripts
05:21 Fake sponsor
07:11 SeaKR: Self-aware Knowledge Retrieval for Adaptive Retrieval Augmented Generation
08:47 Simulating Classroom Education with LLM-Empowered Agents
10:16 Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
12:31 Outro
OpenAI's advanced Voice Mode for ChatGPT Plus users has been delayed, but the company is taking a cautious approach to ensure safety and reliability.
ESM3 is a language model that can simulate 500 million years of evolution, making biology programmable and opening up possibilities for medicine, biology research, and clean energy.
R2R is an open-source project on GitHub that offers a comprehensive and state-of-the-art retrieval-augmented generation system for developers, making it accessible to anyone who wants to try it out.
MG-LLaVA is a new multi-modal large language model that enhances visual processing capabilities by incorporating a multi-granularity vision flow, including low-resolution, high-resolution, and object-centric features.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:36 OpenAI Delays ChatGPT Voice Mode
03:27 ESM3 Simulating 500 million years of evolution with a language model
04:38 Rag to Riches
06:00 Fake sponsor
08:11 MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
09:49 Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon
11:13 Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
13:02 Outro
From the publisher's feed