
Sign up to save your podcasts
Or


Google merges Android, Chrome, and hardware divisions to deliver higher quality products and experiences for users and partners, with a focus on AI innovation.
Boston Dynamics introduces the electric Atlas robot, designed for real-world applications and stronger, more dexterous, and more agile than its predecessors.
"Towards Large Language Models as Copilots for Theorem Proving in Lean" explores using large language models to assist humans in theorem proving.
"AutoCrawler: A Progressive Understanding Web Agent for Web Crawler Generation" introduces AutoCrawler, a framework for generating web crawlers that leverages the power of large language models to handle diverse and changing web environments more efficiently.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:32 Google merges the Android, Chrome, and hardware divisions
03:02 New Atlas Robot from Boston Dynamics
05:01 Karpathi On Llama3
06:19 Fake sponsor
08:14 Towards Large Language Models as Copilots for Theorem Proving in Lean
09:47 AutoCrawler: A Progressive Understanding Web Agent for Web Crawler Generation
11:21 Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
12:58 Outro
Meta announces the release of Llama 3, their new open-source language model with improved reasoning and instruction-following capabilities.
Microsoft invests $1.5 billion in UAE-based AI firm G42, with concerns over its China links requiring negotiations with the Biden administration.
Researchers present "Dynamic Typography," an automated text animation scheme that combines deforming letters to convey semantic meaning and infusing them with movement based on user prompts.
The AI Safety Benchmark from MLCommons is a tool to assess the safety risks of AI systems that use chat-tuned language models, covering 7 of the 13 hazard categories identified by the working group.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:47 Meta Announces Llama 3
03:11 Microsoft invests $1.5B in UAE AI firm
04:59 Randar: A Minecraft exploit that uses LLL lattice reduction to crack server RNG
06:23 Fake sponsor
08:08 Dynamic Typography: Bringing Text to Life via Video Diffusion Prior
09:38 Introducing v0.5 of the AI Safety Benchmark from MLCommons
11:15 BLINK: Multimodal Large Language Models Can See but Not Perceive
13:05 Outro
Boston Dynamics has revealed their new Atlas robot, which boasts impressive dexterity and agility, and is designed for real-world applications.
Stable Assistant, a chatbot powered by Stability AI's text and image generation technology, is now available via an API on the Stability AI developer platform, and features Stable Diffusion 3 and Stable LM 2 12B.
Google DeepMind's "Many-Shot In-Context Learning" proposes a new method of learning from a few examples in a specific context, and found that Reinforced and Unsupervised ICL settings can be quite effective in the many-shot regime.
AWS AI Labs' "Fewer Truncations Improve Language Modeling" introduces a new method called Best-fit Packing that packs documents into training sequences through length-aware combinatorial optimization, and achieved superior performance compared to concatenation.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:25 Boston Dynamics reveals the new Atlas robot
02:47 Stable Diffusion 3 API now available as Stable Assistant effort looms
04:54 Cyc: history's forgotten AI project
06:12 Fake sponsor
08:02 Many-Shot In-Context Learning
09:59 Fewer Truncations Improve Language Modeling
11:38 Can Language Models Solve Olympiad Programming?
13:12 Outro
Adobe is introducing new AI-powered tools to their video editing software, including the ability to extend video clips, add or remove objects from scenes, and generate B-roll footage using prompts.
Amazon's Bedrock platform is adding all three versions of Anthropic's Claude 3 AI model, enhancing the ability of customers to rapidly test, build, and deploy generative AI applications across their organizations.
"The Illusion of State in State-Space Models" challenges the assumption that SSMs are inherently better at state tracking than transformers.
"Megalodon" proposes a new neural architecture for efficient sequence modeling, allowing for unlimited context length and better efficiency than Transformers.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:35 Adobe previews AI video features
02:56 Amazon Puts All Three Claude AI Models on Bedrock
05:07 Automating Complex Business Workflows with Cohere: Multi-Step Tool Use in Action
07:02 Fake sponsor
09:02 The Illusion of State in State-Space Models
10:58 Generative Information Retrieval Evaluation
12:49 Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
14:26 Outro
Reka Core, a comprehensive multimodal solution, is one of only two commercially available models that can handle input from text, images, videos, and audio.
OpenAI's Batch API promises to save costs and increase rate limits on certain async tasks like summarization, translation, and image classification.
COCONut is the largest and most comprehensive segmentation dataset to date, with high-quality annotations and harmonized segmentation types.
DR-PO algorithm directly resets the policy optimizer to the states in the offline dataset, leading to better generative models that are fine-tuned to human preferences.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:52 Reka Core: Our Frontier Class Multimodal Language Model
03:43 OpenAI Batch API
05:14 OpenAI fires two researchers for leaking info
06:42 Fake sponsor
08:54 COCONut: Modernizing COCO Segmentation
10:27 Dataset Reset Policy Optimization for RLHF
12:22 Probing the 3D Awareness of Visual Foundation Models
13:58 Outro
Meta is testing an AI-powered search bar in Instagram, which could improve the quality of search and help users discover new content on the platform.
Grok-1.5V is a new multimodal model that can process a wide variety of visual information and outperforms its peers in the new RealWorldQA benchmark.
"Scaling (Down) CLIP" explores the performance of the Contrastive Language-Image Pre-training (CLIP) when scaled down to limited computation budgets, and shows that smaller datasets and models can still achieve comparable performance.
"Pre-training Small Base LMs with Fewer Tokens" investigates a simple approach called Inheritune to develop a small base language model (LM) from a larger existing LM, which can effectively match the val loss of their bigger counterparts when trained from scratch for the same number of training steps.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:40 Meta is testing an AI-powered search bar in Instagram
03:02 Grok-1.5 Vision Preview
04:56 Visualizing Attention, a Transformer's Heart
06:12 Fake sponsor
08:27 Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies
10:11 Pre-training Small Base LMs with Fewer Tokens
11:58 Flying with Photons: Rendering Novel Views of Propagating Light
13:57 Outro
The Ai Pin, a new device that offloads smartphone tasks, is discussed, funded by OpenAI's Sam Altman and other companies.
A Twitter thread about OpenAI's spider problem is shared, raising questions about the consequences of AI technology.
The paper "Adapting LLaMA Decoder to Vision Transformer" explores adapting decoder-only Transformers to computer vision, resulting in the creation of iLLaMA.
The paper "Exploring Concept Depth" studies how large language models acquire knowledge at different depths, with implications for understanding learning processes and designing models.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:33 This Artificially Intelligent Pin Wants to Free You From Your Phone
03:32 Anyone got a contact at OpenAI. They have a spider problem.
04:47 STORM: Synthesis of Topic Outlines through Retrieval and Multi-perspective Question Asking
06:28 Fake sponsor
08:23 Adapting LLaMA Decoder to Vision Transformer
10:11 RULER: What's the Real Context Size of Your Long-Context Language Models?
12:05 Exploring Concept Depth: How Large Language Models Acquire Knowledge at Different Layers?
13:42 Outro
Meta's Training and Inference Accelerator promises significant performance improvements for AI workloads.
Avi Wigderson receives the Turing Award for his contributions to the theory of computation and randomness in computation.
Intel's Meteor Lake iGPU and Mistral 8x22B offer exciting advancements in the GPU market and language models.
MuPT and Eagle and Finch present new models for music generation and sequence modeling, respectively.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:37 Our next-generation Meta Training and Inference Accelerator
03:02 ACM A.M. Turing Award Honors Avi Wigderson for Foundational Contributions to the Theory of Computation
05:05 Intel’s Ambitious Meteor Lake iGPU
06:06 Mistral 8x22B
07:34 Fake sponsor
09:34 MuPT: A Generative Symbolic Music Pretrained Transformer
11:11 Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence
12:49 Outro
Meta is launching its Llama 3 open source LLM with 140 billion parameters, catching up to OpenAI's ChatGPT.
Intel's Gaudi 3 AI accelerator is breaking down proprietary walls to bring choice to enterprise GenAI market, with OEMs like Dell and Lenovo adopting it.
Microsoft Research's Direct Nash Optimization (DNO) is a new approach to improving language models, achieving state-of-the-art win-rates against GPT-4-Turbo.
UniFL is a unified framework that uses feedback learning to enhance diffusion models, improving both the quality of generated models and their acceleration.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:47 Meta confirms that its Llama 3 open source LLM is coming in the next month
03:17 Intel Breaks Down Proprietary Walls to Bring Choice to Enterprise GenAI Market
05:18 QCon London: Meta Used Monolithic Architecture to Ship Threads in Only Five Months
06:42 Fake sponsor
09:10 Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
10:48 UniFL: Improve Stable Diffusion via Unified Feedback Learning
12:24 MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
14:13 Outro
Sam Altman and Jony Ive are raising funding for a secret AI device company that could challenge the conventional smartphone experience and explore new interaction modalities with artificial intelligence. Tesla is unveiling its new 'robotaxi' on August 8th, which is specifically designed for ridesharing and could potentially shift the company's focus away from achieving self-driving on their existing fleet. "Stream of Search (SoS): Learning to Search in Language" is a paper that proposes a new approach to teach language models how to search by representing the search process in language, as a flattened string. "AutoWebGLM: Bootstrap and Reinforce a Large Language Model-based Web Navigating Agent" is a paper that introduces an automated web navigation agent that uses a Large Language Model (LLM) to browse the web, which outperforms the state-of-the-art chatbot-based and rule-based methods.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:30 Sam Altman and Jony Ive Raising Funding for Secret AI Device Company
02:58 Tesla is unveiling its new ‘robotaxi’ on August 8
04:27 llm.c by Andrej Karpathy
05:43 Fake sponsor
07:41 Stream of Search (SoS): Learning to Search in Language
09:11 Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model
10:51 AutoWebGLM: Bootstrap And Reinforce A Large Language Model-based Web Navigating Agent
12:52 Outro
From the publisher's feed