
Sign up to save your podcasts
Or


This episode dives into Nvidia's stock struggles amid rising competition, while also unpacking Meta's AI blunders and the implications of "hallucinations" in tech. We explore cutting-edge superconducting microprocessors that promise unprecedented energy efficiency and highlight groundbreaking AI research, including eavesdropping techniques and advancements in reinforcement learning.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:50 Nvidia Sank Again Today -- Time to Buy the Artificial Intelligence (AI) Growth Stock Hand Over Fist?
03:09 Meta blames hallucinations after its AI said Trump rally shooting didn’t happen
04:52 Superconducting Microprocessors? Turns Out They're Ultra-Efficient
06:07 Fake sponsor
07:48 Deep-TEMPEST: Using Deep Learning to Eavesdrop on HDMI from its Unintended Electromagnetic Emanations
09:22 SAPG: Split and Aggregate Policy Gradients
10:45 MindSearch: Mimicking Human Minds Elicits Deep AI Searcher
12:44 Outro
This episode dives into Google’s Gemma 2, which claims to outperform GPT-3.5 while tackling responsible AI practices. We explore Black Forest Labs' Flux model, featuring 12 billion parameters and tailored versions for various users. Olivia sheds light on the ethical concerns surrounding the resurgence of pseudoscience in machine learning, particularly physiognomy. Lastly, Belinda reviews critical research on AI safety, advocating for clearer metrics to prevent misleading claims about safety advancements.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:37 Google’s tiny AI model bests GPT-3.5
02:48 Announcing Flux by Black Forest Labs: The Next Leap in Text-to-Image Models
04:28 The reanimation of pseudoscience in machine learning and its ethical repercussions
06:06 Fake sponsor
08:04 MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
09:55 Spectra: A Comprehensive Study of Ternary, Quantized, and FP16 Language Models
11:41 Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
13:33 Outro
OpenAI's new prototype, SearchGPT, promises to combine AI smarts with real-time web information to make search easier.
AI has achieved silver-medal standards at the International Mathematical Olympiad, raising questions about the future of mathematics and the role of AI in solving complex problems.
The reliability of AI existential risk probabilities is called into question in a thought-provoking article, challenging the authority we often assign to these forecasts and calling for more scrutiny.
Three fascinating papers from UNC Chapel Hill, Google DeepMind, and a collaboration between Caltech and NVIDIA explore advancements in theorem proving, balancing fast and slow planning, and aligning large language models with Best-of-N distillation. These papers could transform the way we approach complex problems with language models and streamline the development of LLMs.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:54 OpenAI Announces SearchGPT
03:15 AI achieves silver-medal standard solving International Mathematical Olympiad problems
04:55 AI existential risk probabilities are too unreliable to inform policy
06:25 Fake sponsor
08:21 LeanDojo: Theorem Proving with Retrieval-Augmented Language Models
10:10 System-1.x: Learning to Balance Fast and Slow Planning with Language Models
12:01 BOND: Aligning LLMs with Best-of-N Distillation
13:43 Outro
Mistral Large 2 release with advanced features and multilingual support.
Elon Musk's announcement of the Memphis Supercluster for creating the world's most powerful AI.
Discussion of emergence in complex systems and the MINT-1T dataset for training large multimodal models.
Introduction of OpenDevin, an open platform for developing AI agents and MOMAland, a benchmark framework for multi-objective multi-agent reinforcement learning.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:39 Mistral Large 2 Release
03:01 Elon Musk Announces Memphis Supercomputer
04:48 The Puzzle of How Large-Scale Order Emerges in Complex Systems
06:22 Fake sponsor
08:37 MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
10:16 OpenDevin: An Open Platform for AI Software Developers as Generalist Agents
11:53 MOMAland: A Set of Benchmarks for Multi-Objective Multi-Agent Reinforcement Learning
13:31 Outro
This episode features the introduction of Llama 3.1, Meta's cutting-edge AI model with remarkable flexibility and extensive language support. We delve into Alphabet's impressive 14% revenue growth, highlighting the increasing demand for AI infrastructure in cloud computing. The System-1.x Planner is explored, demonstrating its innovative balance between fast and slow planning modes, leading to enhanced performance. Finally, we discuss MovieDreamer, a groundbreaking model that elevates video generation by ensuring narrative coherence and high visual quality in long-form content.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:39 Introducing Llama 3.1: Our most capable models to date
02:59 Alphabet revenue jump shows no sign of AI denting search business
04:36 Open Source AI Is the Path Forward
05:40 Fake sponsor
07:41 System-1.x: Learning to Balance Fast and Slow Planning with Language Models
09:31 KAN or MLP: A Fairer Comparison
11:08 MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequence
12:53 Outro
Meta's upcoming Llama 3.1 models could outperform the current state-of-the-art closed-source LLM model, OpenAI's GPT-4o.
OpenAI is planning to develop its own AI chip to optimize performance and potentially supercharge their progress towards AGI.
Apple's SlowFast-LLaVA is a new training-free video large language model that captures both detailed spatial semantics and long-range temporal context in video without exceeding the token budget of commonly used LLMs.
Google's Conditioned Language Policy (CLP) framework is a general framework that builds on techniques from multi-task training and parameter-efficient finetuning to develop steerable models that can trade-off multiple conflicting objectives at inference time.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:28 LLAMA 405B Performance Leaked
03:01 OpenAI Wants Its Own AI Chips
04:25 Towards more cooperative AI safety strategies
06:01 Fake sponsor
07:35 SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
09:17 AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
10:56 Conditioned Language Policy: A General Framework for Steerable Multi-Objective Finetuning
12:46 Outro
Claude for Android is now available, bringing AI-powered assistance to a wider audience.
MIT researchers have developed a new machine-learning framework that can predict materials' thermal properties up to 1,000 times faster than other AI-based techniques, potentially improving energy efficiency.
TinkerBird, a vector database designed for efficient storage and retrieval of high-dimensional vectors, is disrupting traditional RAG workflows and eliminating roundtrip delays associated with client-server models.
ChatQA 2, a Llama3-based model from NVIDIA, bridges the gap between open-access LLMs and leading proprietary models in long-context understanding and retrieval-augmented generation capabilities, while Stable Audio Open, an open-weights text-to-audio model from Stability AI, showcases potential for high-quality stereo sound synthesis at 44.1kHz.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:34 Claude for Android is here
02:50 AI method radically speeds predictions of materials’ thermal properties
04:44 TinkerBird
06:10 Fake sponsor
08:10 ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities
09:54 Stable Audio Open
11:28 Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
13:54 Outro
OpenAI has released their newest model, GPT-4o mini, which is more cost-efficient and excels in mathematical reasoning and coding tasks.
NVIDIA's Mistral NeMo 12B is a state-of-the-art language model with unprecedented accuracy and enterprise-grade support.
A new speech recognition keyboard and service for Android called Transcribro has been developed, which is private and on-device.
Research papers explore the impact of vocabulary size on language model scaling, the use of large datastores for retrieval-based language models, and a method for generating long sequences of views of a cityscape using AI and computer vision.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:40 OpenAI Announces GPT 4o mini
03:11 Mistral AI and NVIDIA Unveil Mistral NeMo 12B, a Cutting-Edge Enterprise AI Model
05:28 Transcribro: Private and on-device speech recognition keyboard and service for Android
06:43 Fake sponsor
08:49 Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies
10:19 Scaling Retrieval-Based Language Models with a Trillion-Token Datastore
11:49 Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffusion
13:26 Outro
Apple, Nvidia, Anthropic, and Salesforce caught using content without creators' consent for AI training.
Mistral AI launches two new open-source models, Codestral Mamba and Mathstral, with impressive capabilities.
NVIDIA transitions to fully open-source GPU kernel modules, offering new capabilities and easy switching for users.
Exciting research papers include Ref-AVS for multimodal object segmentation, Qwen2-Audio for large-scale audio-language modeling, and DiT-MoE for scalable language modeling and image generation.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:27 Apple, Nvidia, Anthropic Used Thousands of Swiped YouTube Videos to Train AI
02:46 Mistral's New Open Source Models
04:09 NVIDIA Transitions Fully Towards Open-Source GPU Kernel Modules
05:37 Fake sponsor
07:15 Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
08:47 Qwen2-Audio Technical Report
10:49 Scaling Diffusion Transformers to 16 Billion Parameters
12:21 Outro
From the publisher's feed