
Sign up to save your podcasts
Or


Soundry AI promises to be a game-changer for musicians with its superior flexibility and versatility compared to standard sample libraries.
OpenAI is set to release GPT-5, an improved version of the AI language model that powers ChatGPT, which could represent a notable advancement for OpenAI.
RewardBench, a benchmark dataset and codebase for evaluating reward models, provides a standardized way to evaluate reward models on a range of tasks, including chat, reasoning, and safety.
DepthFM's Fast Monocular Depth Estimation with Flow Matching is a promising direction for the field of monocular depth estimation, with its generative approach and state-of-the-art performance on standard benchmarks.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:33 Soundry, AI for Musicians, by Musicians
02:55 GPT-5 might arrive this summer as a “materially better” update to ChatGPT
04:55 The Reddits
06:38 Fake sponsor
08:33 RewardBench: Evaluating Reward Models for Language Modeling
10:11 Evaluating Frontier Models for Dangerous Capabilities
11:25 DepthFM: Fast Monocular Depth Estimation with Flow Matching
12:57 Outro
Microsoft hires Mustafa Suleyman to lead AI division, including Copilot, an AI tool that helps programmers write better code.
NVIDIA announces Project GR00T, a multimodal AI system that enables advanced humanoid robots to learn skills and interact with the real world, with partnerships from leading robotics companies.
LLMLingua-2 and Agent-FLAN are two cutting-edge AI research papers that address the challenges of prompt compression and incorporating agent ability into large language models, respectively.
uBlockOrigin and uBlacklist's huge AI blocklist is a useful resource for those who want to eradicate AI-generated content from their search results.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:52 Microsoft hires Suleyman to lead AI
03:08 NVIDIA Announces Project GR00T Foundation Model for Humanoid Robots and Major Isaac Robotics Platform Update
04:45 uBlockOrigin & uBlacklist Huge AI Blocklist
05:55 Fake sponsor
07:50 LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression
09:22 mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
11:14 Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models
13:17 Outro
Saudi Arabia's $40 billion investment in AI, making them the largest investor in the field.
OpenAI's CEO Sam Altman's prediction that compute will be the currency of the future.
Vid2Robot, a video-based learning framework that directly produces robot actions given a video demonstration of a manipulation task and the robot's current visual observations.
TexDreamer, the first zero-shot multimodal high-fidelity 3D human texture generation model.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:50 Saudi Arabia Plans $40 Billion Push Into Artificial Intelligence
03:14 Sam Altman on GPT-5 and the future of AI
04:46 Price Per Part in Lego Sets
05:57 Fake sponsor
07:52 Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers
09:35 mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
11:31 TexDreamer: Towards Zero-Shot High-Fidelity 3D Human Texture Generation
13:20 Outro
Nvidia's Blackwell B200 GPU is the world's most powerful chip for AI, offering up to 20 petaflops of FP4 horsepower and reducing cost and energy consumption.
Stable Video 3D is a generative model that can be used for commercial purposes and improves the quality of 3D meshes generated directly from novel views.
Apple is in talks with potential partners like Google to integrate their AI models into the iPhone and the broader OS ecosystem, which could be a major win for both companies.
The latest AI research includes papers on the security of API-protected LLMs, optimizing attention modules for processing long sequences, and super-resolution of anime images.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:40 Nvidia reveals Blackwell B200 GPU, the ‘world’s most powerful chip’ for AI
03:46 Introducing Stable Video 3D: Quality Novel View Synthesis and 3D Generation from Single Images
05:39 Apple's AI plans reportedly could involve a partner like Google
06:51 Fake sponsor
08:52 Logits of API-Protected LLMs Leak Proprietary Information
10:35 BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
12:19 APISR: Anime Production Inspired Real-World Anime Super-Resolution
13:59 Outro
Figure, a leading AI robotics company, is making significant advancements in creating robots that can perceive their environment, make decisions, and take action, all in a way that aligns with human expectations.
OpenAI may have accidentally leaked details about a new AI model called GPT-4.5 Turbo, which could level the playing field with Google's AI model Gemini.
Two papers explore the development and evaluation of large language models (LLMs) for code-related tasks, and propose simple and scalable strategies to continually pre-train LLMs to save on compute.
Another paper investigates scaling in the over-trained regime and relates language model perplexity to downstream task performance via a power law, providing useful insights into how language models can be scaled and evaluated more effectively.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:28 Figure AI
03:02 Did OpenAI just accidentally leak the next big ChatGPT upgrade?
04:48 Gradio's Grog
05:47 Fake sponsor
07:37 Simple and Scalable Strategies to Continually Pre-train Large Language Models
09:31 LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
11:11 Language models scale reliably with over-training and on downstream tasks
13:07 Outro
X AI has released their 314 billion parameter Mixture-of-Experts model, Grok-1, which is currently the largest language model that has been publicly released and could be used for a variety of tasks.
The European Parliament has passed a landmark AI act that bans certain AI applications and requires strict obligations for high-risk AI, positioning itself as the global standard for regulation.
The paper "MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training" explores the importance of Multimodal Large Language Models (MLLM) and how a careful mix of image-caption, interleaved image-text, and text-only data is crucial for achieving state-of-the-art few-shot results across multiple benchmarks.
The paper "Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews" examines the impact of large language models on scientific peer review and suggests that LLM-generated text can affect the quality and fairness of peer review.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:24 Open Release of Grok-1
02:43 Claude 3 Haiku
04:34 EU passes landmark AI act
06:15 The Rest of the World Disappears’: Claire Voisin on Mathematical Creativity
07:59 Fake sponsor
09:54 MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
11:30 Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
13:19 Outro
Introducing Devin, the first AI software engineer that can plan and execute complex engineering tasks requiring thousands of decisions.
Google's AI chatbot won't answer questions about upcoming elections to prevent inaccurate or misleading responses.
WorkArena, a benchmark measuring the ability of large language model-based agents to perform tasks that align with the daily work of knowledge workers using enterprise software systems.
Synth$^2$, a novel approach that leverages Large Language Models (LLMs) and image generation models to create synthetic image-text pairs for efficient and effective Visual-Language Model (VLM) training.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:59 Introducing Devin, the first AI software engineer
03:48 Google won’t let you use its Gemini AI to answer questions about an upcoming election in your country
05:35 AI Datacenter Energy Dilemma - Race for AI Datacenter Space
06:45 Fake sponsor
09:08 Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
10:26 WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
11:52 Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings
13:38 Outro
The latest research in AI language models, including algorithmic progress and multistep consistency models.
A new LLM model called Command-R, designed for large-scale production workloads.
The announcement of Meta's investment in AI infrastructure, including two 24k GPU clusters and plans for continued growth.
A discussion on the potential security risks of model-stealing attacks on language models.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:42 Command-R: Retrieval Augmented Generation at Production Scale
03:48 Building Meta’s GenAI Infrastructure
05:41 MLX Server
06:37 Fake sponsor
08:45 Algorithmic progress in language models
10:23 Multistep Consistency Models
11:35 Stealing Part of a Production Language Model
13:31 Outro
Nvidia's CEO claims that even free AI chips from competitors can't beat Nvidia's GPUs, highlighting the company's dominance in the AI industry.
The energy demands of AI and cryptocurrency are discussed, with concerns raised about the potential consequences of their growing electricity needs.
OpenAI's Superalignment team has developed a powerful new tool called Transformer Debugger, which allows researchers to investigate specific behaviors of small language models.
Three AI research papers are discussed, including a model-stealing attack on black-box production language models, a study questioning the effectiveness of cosine-similarity in measuring semantic similarity, and the latest version of SPLADE, a library for ranking natural language queries in information retrieval systems.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:49 Jensen Huang says even free AI chips from his competitors can't beat Nvidia's GPUs
03:42 The Obscene Energy Demands of A.I.
05:03 Transformer Debugger
06:27 Fake sponsor
08:07 Stealing Part of a Production Language Model
09:37 Is Cosine-Similarity of Embeddings Really About Similarity?
11:05 SPLADE-v3: New baselines for SPLADE
12:46 Outro
Apple's upcoming AI releases in 2024, including a new Siri upgrade and a rumored "AppleGPT" LLM.
Sam Altman's return to OpenAI's board of directors after an investigation into his ouster.
Research on high-resolution image synthesis using rectified flow transformers and a novel transformer-based architecture for text-to-image generation.
A new metric called FENICE for evaluating factual inconsistencies in automatically generated text summaries.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:45 Apple Preparing AI Releases for 2024
03:03 Sam Altman rejoins OpenAI board of directors as investigation into his ouster comes to a close
05:11 Your guide to AI: March 2024
06:45 Fake sponsor
08:41 Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
10:50 FENICE: Factuality Evaluation of summarization based on Natural language Inference and Claim Extraction
12:19 Poly-View Contrastive Learning
14:16 Outro
From the publisher's feed