
Sign up to save your podcasts
Or


A Google engineer has been indicted for allegedly stealing over 500 confidential files containing AI trade secrets while working for China-based companies seeking an edge in the AI technology race.
A tutorial series explores parallelism strategies for training large deep learning models, making it accessible to everyone regardless of the hardware you have available.
Value functions are a crucial component in deep reinforcement learning, and a new approach using categorical cross-entropy instead of regression can significantly improve performance and scalability in a variety of domains.
Backtracing is the task of retrieving the text segment that most likely caused a user query, and it can help improve content delivery and communication by identifying linguistic triggers that influence user queries.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:33 Google engineer indicted over allegedly stealing AI trade secrets for China
03:57 Training Models at Scale Tutorial
05:24 Autogenerating a Book Series From Three Years of iMessages
06:22 Fake sponsor
08:16 Design2Code: How Far Are We From Automating Front-End Engineering?
10:09 Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
11:43 Backtracing: Retrieving the Cause of the Query
13:27 Outro
Perplexity AI is a search startup that's looking to take on Google by solving the inadequacies of searching the web. They are nearing unicorn status with a valuation of around $1 billion.
Microsoft is being sued by The New York Times for copyright infringement and abusing the newspaper’s intellectual property in training LLMs. Microsoft accuses the Times of "unsubstantiated" claims and compares the lawsuit to Hollywood's resistance to the VCR in the 70s.
A new paper introduces the concept of General Computer Control (GCC), which is the idea of building agents that can master any computer task by taking only screen images and producing keyboard and mouse operations as output. The authors propose a framework called Cradle that has strong reasoning abilities to ensure generalizability and self-improvement across various tasks.
A paper evaluates different tokenizer inference methods and their impact on the performance of downstream NLP tasks. The authors found that for the most commonly used tokenizers, greedy inference performs surprisingly well, and a recently-introduced contextually-informed tokenizer outperforms all others on morphological alignment.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:23 Perplexity Poised To Become Latest AI Startup To Hit Unicorn Status — Report
02:53 Microsoft compares The New York Times’ claims against OpenAI to Hollywood’s early fight against VCR
04:41 Training great LLMs entirely from ground zero in the wilderness as a startup
05:49 Fake sponsor
07:39 Towards General Computer Control: A Multimodal Agent for Red Dead Redemption II as a Case Study
09:15 Design2Code: How Far Are We From Automating Front-End Engineering?
11:01 Greed is All You Need: An Evaluation of Tokenizer Inference Methods
12:49 Outro
Groq, an AI chip startup, forms a new business unit and acquires Definitive Intelligence to expand its customer and developer ecosystem.
OpenAI responds to Elon Musk's lawsuit, revealing that Musk himself wanted "absolute control" over the company by merging it with Tesla.
A new Postgres extension called pg_vectorize automates the transformation and orchestration of text to embeddings, providing workflows for vector search and RAG.
UNITS, a unified time series model, achieves superior performance compared to task-specific models and repurposed natural language-based LLMs, demonstrating remarkable zero-shot, few-shot, and prompt learning capabilities.
Contact: [email protected]
Timestamps:
00:34 Introduction
02:05 AI chip startup Groq forms new business unit, acquires Definitive Intelligence
03:49 OpenAI says Elon Musk wanted ‘absolute control’ of the company
05:29 pg_vectorize: a VectorDB for Postgres
06:29 Fake sponsor
08:26 Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
09:54 UniTS: Building a Unified Time Series Model
11:39 DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving
13:14 Outro
Anthropic's new and improved Claude 3 model family sets new industry benchmarks across a wide range of cognitive tasks, exhibiting near-human levels of comprehension and fluency on complex tasks.
Elon Musk is suing OpenAI and CEO Sam Altman for allegedly abandoning their original mission to benefit humanity and instead focusing on profits with Microsoft.
Opus 1.5 brings quality improvements, including machine learning-based upgrades, while remaining fully compatible with RFC 6716, and uses deep learning techniques to process or generate signals themselves.
The Multimodal ArXiv dataset represents an important step forward for LVLMs when it comes to interpreting and understanding complex scientific figures, achieving a 10.4% absolute accuracy gain on a multimodal mathematical reasoning benchmark.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:27 Introducing the next generation Claude: Claude 3
03:17 Elon Musk sues Sam Altman and OpenAI
04:59 Opus Gets a Serious Machine Learning Upgrade
06:29 Fake sponsor
08:28 UniTS: Building a Unified Time Series Model
10:10 Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
12:06 Learning and Leveraging World Models in Visual Representation Learning
13:48 Outro
Adobe's new generative AI tools for custom audio creation and editing.
Tumblr and WordPress selling user data to train AI tools, sparking backlash.
MOSAIC, a modular system for assistive and interactive cooking using natural language and multiple robots.
A new approach to real-world humanoid control using a causal transformer model trained through autoregressive prediction of sensorimotor trajectories.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:20 Adobe previews new cutting-edge generative AI tools for crafting and editing custom audio
02:39 Tumblr and WordPress to Sell Users’ Data to Train AI Tools
04:14 “AI will cure cancer” misunderstands both AI and medicine
05:56 Fake sponsor
08:12 MOSAIC: A Modular System for Assistive and Interactive Cooking
09:54 Humanoid Locomotion as Next Token Prediction
11:50 In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
13:22 Outro
Meta Platforms is set to launch its new AI language model, Llama 3, which promises to tackle taboo questions with more grace and respect than its predecessor.
Apple is ramping up its investment in GenAI, with plans to upgrade Siri and iOS’ built-in search tool, Spotlight, with GenAI models to handle more complex queries and multi-turn conversations.
The University of California, Berkeley, has published a paper exploring unsupervised zero-shot reinforcement learning via functional reward encodings, which could enable pre-training of an agent to adapt to any new downstream tasks in a zero-shot manner.
TrustMol, an inverse molecular design method built to be trustworthy, has been proposed by the Max Planck Institute for Informatics, which could make the IMD process more explainable and reliable.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:52 Meta plans launch of new AI language model Llama 3 in July, The Information reports
02:56 Tim Cook says Apple will ‘break new ground’ in GenAI this year
04:35 Things You Should Never Do, Part I
05:46 Fake sponsor
07:28 Unsupervised Zero-Shot Reinforcement Learning via Functional Reward Encodings
09:04 TrustMol: Trustworthy Inverse Molecular Design via Alignment with Molecular Dynamics
10:57 Stochastic Gradient Succeeds for Bandits
13:23 Outro
Google's image creation tool, Gemini, has been generating offensive and embarrassing results, prompting the company to make structural changes and update product guidelines to avoid bias in AI tools.
C3.ai, a software maker that helps companies build AI applications, reported a narrower-than-expected loss and revenue that topped estimates, causing AI stock to pop more than 14% in extended trading.
A new paper introduces a cost-effective Large Language Model called a 1-bit LLM, which matches the performance of full-precision Transformer LLMs while being significantly more efficient in terms of latency, memory, throughput, and energy consumption.
Another paper proposes a hybrid approach that combines a frozen LLM with a small language model to improve the efficiency of autoregressive decoding for Large Language Models, resulting in substantial speedups of up to 4 times with minor performance penalties. Additionally, a new framework called EMO utilizes a direct audio-to-video synthesis approach to produce highly expressive and lifelike talking head videos.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:38 Google CEO calls AI tool’s controversial responses ‘completely unacceptable’
03:11 Artificial Intelligence Play C3.ai Climbs On Earnings Report, Outlook
04:41 Jason Wei On Sora
06:19 Fake sponsor
08:35 The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
09:19 Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding
10:38 EMO: Emote Portrait Alive - Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions
12:04 Outro
DeepMind's Genie, a tool that creates video games with just a prompt or an image, is a game-changer in the industry.
Tyler Perry's $800M studio expansion is on hold after seeing OpenAI's Sora, highlighting the potential for AI to replace human workers in the entertainment industry.
MobileLLM is a promising development for those looking to deploy efficient language models on mobile devices.
A comprehensive review of existing literature on data selection methods for language models provides a taxonomy of existing approaches and proposes promising avenues for future research.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:44 DeepMind's Genie: creating videogames with prompts
02:48 Tyler Perry Puts $800M Studio Expansion on Hold After Seeing OpenAI’s Sora: “Jobs Are Going to Be Lost”
04:41 Speakz AI
06:15 Fake sponsor
07:56 MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
09:57 A Survey on Data Selection for Language Models
11:26 Do Large Language Models Latently Perform Multi-Hop Reasoning?
13:44 Outro
Mistral AI has launched a new conversational assistant, Le Chat Mistral, which serves as an entry point to interact with their various models. They're also launching Le Chat Enterprise, which could be useful for businesses looking to boost productivity and efficiency.
Microsoft has partnered with Mistral, a French company focused on language models, and will be taking a minor stake in the company and offering their language models on Azure AI platform. Mistral is also releasing a new model called Mistral Large, which is designed to compete with OpenAI's GPT-4 model.
"Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models" by Levy et al. investigates how the performance of Large Language Models (LLMs) changes when the input length is extended. The authors found that there is a notable degradation in LLMs' reasoning performance at much shorter input lengths than their technical maximum.
"Executable Code Actions Elicit Better LLM Agents" proposes using executable Python code to consolidate LLM agents' actions into a unified action space called CodeAct. CodeAct outperforms widely used alternatives by up to 20% higher success rate and could have a lot of practical applications.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:39 Le Chat announced by Mistral AI
02:53 Microsoft partners with Mistral in second AI deal beyond OpenAI
04:29 Introducing Phind 70Billion
05:27 Fake sponsor
07:05 Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
08:47 Executable Code Actions Elicit Better LLM Agents
10:36 Cleaner Pretraining Corpus Curation with Neural Web Scraping
12:20 Outro
Gemma, a new family of lightweight, state-of-the-art open models built for responsible AI development, is introduced by Google.
"Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models" presents a new method for instruction tuning of Large Language Models (LLMs) called Generalized Instruction Tuning (GLAN).
"MuLan: Multimodal-LLM Agent for Progressive Multi-Object Diffusion" addresses the challenge of generating images of multiple objects with spatial relationships and attribute bindings.
"Instruction-tuned Language Models are Better Knowledge Learners" explores how to update factual knowledge in large language models.
Contact: [email protected]
Timestamps:
00:34 Introduction
01:21 Google DeepMind Releases Gemma
03:28 Andrej Karpathy on Gemma's Tokenizer
04:16 Groq Inference Tokenomics: Speed, But At What Cost?
05:51 Fake sponsor
07:44 Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models
09:38 MuLan: Multimodal-LLM Agent for Progressive Multi-Object Diffusion
11:06 Instruction-tuned Language Models are Better Knowledge Learners
12:58 Outro
From the publisher's feed