GPT Reviews

GPT Reviews

By EarkindNewsDaily News
Download on the App Store

GPT Reviews episodes

  • Andrej Karpathi starts AI Ed Company 🧑‍🏫 // xLSTM for Time Series Forecasting 📊 // Customized Video Generation with Still-Moving 📹

    AI expert Andrej Karpathy is starting an AI+Education company called Eureka Labs, which aims to create an ideal experience for learning something new by leveraging recent progress in generative AI.

    YouTube Music is rolling out new AI-powered features, including 'Sound Search' that allows users to search YouTube's catalog of over 100 million songs by singing, humming, or playing a tune, and 'AI-generated conversational radio' that creates a tailored playlist based on natural language prompts.

    "xLSTMTime: Long-term Time Series Forecasting with xLSTM" explores the use of extended LSTM (xLSTM) for improving long-term time series forecasting, demonstrating superior forecasting capabilities compared to other state-of-the-art models.

    "Still-Moving: Customized Video Generation without Customized Video Data" introduces a framework called Still-Moving that seamlessly integrates the spatial prior of a customized text-to-image (T2I) model with the motion prior of a T2V model, achieving impressive results on personalized, stylized, and conditional video generation tasks.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:57 Andrej Karpathy starting an AI Education Company

    03:27 YouTube Music sound search rolling out, AI ‘conversational radio’ in testing

    05:12 GB200 Hardware Architecture & Component Supply Chain & BOM

    06:50 Fake sponsor

    08:34 xLSTMTime : Long-term Time Series Forecasting With xLSTM

    10:04 Still-Moving: Customized Video Generation without Customized Video Data

    11:47 NeedleBench: Can LLMs Do Retrieval and Reasoning in 1 Million Context Window?

    13:32 Outro

    15 min
  • OpenAI's Reasoning Secret Project 🤔 // Mysterious Models in LMSys Arena 👀 // Human-like Episodic Memory 🧠

    OpenAI's mysterious code name 'Strawberry' and its potential to revolutionize AI capabilities.

    The appearance of new models in the LMSYS Chatbot Arena, potentially hinting at a new mini-GPT release from OpenAI.

    The introduction of EM-LLM, a new model that integrates human episodic memory and event cognition into LLMs, improving their ability to process extensive contexts.

    The creation of SPIQA, a large-scale question answering dataset specifically designed to interpret complex figures and tables within the context of scientific research articles.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:44 Exclusive: OpenAI working on new reasoning technology under code name ‘Strawberry’

    03:25 Mysterious Models in LMSys Arena

    04:51 Run Cuda on AMD GPUs

    06:14 Fake sponsor

    07:52 Human-like Episodic Memory for Infinite Context LLMs

    09:56 Toto: Time Series Optimized Transformer for Observability

    11:28 SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers

    13:21 Outro

    15 min
  • Gemini 1.5 Pro Navigating The World 🤖 // OpenAI's Progress Tiers 🏆 // AI Memory Aids and Ethics 🧠

    Google DeepMind's latest research on robot navigation using the Gemini 1.5 Pro.

    OpenAI's new five-tier system to track progress towards artificial general intelligence.

    The development of AI-powered memory aids and the ethical implications of relying on technology for memory recall.

    Microsoft's innovative encoding framework, SheetCompressor, for large language models to understand and reason about spreadsheets.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:43 Gemini 1.5 Pro’s long context window help robots navigate the world?

    03:00 OpenAI Scale Ranks Progress Toward ‘Human-Level’ Problem Solving

    04:47 Inside the AI memory machine

    05:56 MambaVision: A Hybrid Mamba-Transformer Vision Backbone

    07:01 Fake sponsor

    09:09 Still-Moving: Customized Video Generation without Customized Video Data

    10:41 SpreadsheetLLM: Encoding Spreadsheets for Large Language Models

    12:51 Outro

    15 min
  • Microsoft and Apple drop OpenAI Board Plans 🤝 // FlashAttention-3 speeds up Attention 🚀 // Improving Mathematical Reasoning 🔢

    Microsoft and Apple drop OpenAI Board plans due to increased regulatory scrutiny in the AI sector.

    Research papers on enhancing mathematical reasoning capabilities of large language models and improving mathematical problem-solving capabilities in visual contexts using Multi-modal Large Language Models (MLLMs).

    FlashAttention-3, an algorithm that speeds up attention mechanism in large language models by up to 2 times faster than previous versions, while maintaining accuracy with lower precision numbers.

    Adaptive In-Context Learning, a technique that simplifies the overall machine learning pipeline, making it more accessible for more organizations.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    02:01 Microsoft, Apple Drop OpenAI Board Plans as Scrutiny Grows

    03:41 Reproducing GPT-2 in C and CUDA

    04:48 Adaptive In-Context Learning

    06:25 FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

    07:53 Fake sponsor

    10:06 Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On

    11:35 MAVIS: Mathematical Visual Instruction Tuning

    13:38 Outro

    15 min
  • Oxygen Initiative 💨 // OpenAI in China 🚫 // AMD's Silo AI Acquisition 💰

    a16z's Oxygen initiative to provide AI startups with access to GPUs at below-market rates in exchange for equity, potentially reshaping the AI VC landscape.

    OpenAI's decision to block users in China from accessing its tools and services, potentially accelerating the development of homegrown models by Chinese AI companies.

    AMD's acquisition of Silo AI for $665M to enhance its AI chip capabilities and compete against industry leader Nvidia.

    Three exciting AI papers on the limitations of vision language models, a new approach for video instruction tuning, and a novel framework for multi-agent collaboration called the Internet of Agents.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:35 a16z is trying to keep AI alive with Oxygen initiative

    03:14 Chinese developers scramble as OpenAI blocks access in China

    04:52 AMD to acquire Finnish startup Silo AI for $665M to step up in AI race

    06:18 Fake sponsor

    08:00 Vision language models are blind

    09:30 Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

    11:17 Internet of Agents: Weaving a Web of Heterogeneous Agents for Collaborative Intelligence

    13:37 Outro

    15 min
  • Multi-Task Language Understanding 📈 // Composable Interventions 🤝 // ARMT Sets Performance Record 💪

    The MNLU-Pro dataset is a more robust and challenging massive multi-task language understanding dataset that's tailored to more rigorously benchmark large language models' capabilities.

    The Composable Interventions framework allows researchers to study the effects of using multiple interventions on a language model, and the order in which interventions are applied can have a significant impact on their effectiveness.

    The MJ-Bench benchmark evaluates the effectiveness of different types of multimodal judges in providing feedback for text-to-image generation models, and the experiments reveal that close-source VLMs generally provide better feedback.

    The Associative Recurrent Memory Transformer (ARMT) is an approach that combines transformer self-attention for local context with segment-level recurrence for storage of task-specific information distributed over a long context, and it sets a new performance record in the recent BABILong multi-task long-context benchmark.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:32 MNLU-Pro Release on HuggingFace Datasets

    03:48 Extrinsic Hallucinations in LLMs

    04:53 RouteLLM

    06:13 Fake sponsor

    08:14 Composable Interventions for Language Models

    09:45 MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?

    11:31 Associative Recurrent Memory Transformer

    13:30 Outro

    15 min
  • OpenAI Breach 🔒 // Google vs. ChatGPT 🔍 // Test-Time Training Layers 👨‍🎓

    OpenAI's internal messaging systems were breached last year, and sensitive details about the company's technology were stolen, raising questions about OpenAI's security protocols and how they handle these types of incidents.

    Despite the popularity of generative AI chatbots like ChatGPT, Google Search's market dominance is actually growing, which may not bode well with antitrust regulators.

    "Learning to (Learn at Test Time): RNNs with Expressive Hidden States" proposes a new class of sequence modeling layers called Test-Time Training (TTT) layers, which have the potential to be a powerful tool for sequence modeling.

    "Finding Visual Task Vectors" introduces a technique called visual prompting, which has potential applications in areas such as robotics and autonomous systems, and provides an alternative to traditional supervised learning.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:38 A Hacker Stole OpenAI Secrets

    03:05 ChatGPT might rule the AI chatbots — but it can't beat Google Search

    04:52 MobileLLM

    06:17 Fake sponsor

    08:20 Learning to (Learn at Test Time): RNNs with Expressive Hidden States

    10:17 Many-Shot In-Context Learning

    11:53 Finding Visual Task Vectors

    13:38 Outro

    15 min
  • Moshi, the new voice-first LM 🗣️ // AI-generated iconic voices 🎙️ // Large language models oversight 🤖

    Moshi, the first real-time AI voice assistant with 70 different emotions and speaking styles, has been unveiled by French startup Kyutai.

    ElevenLabs' Reader App now features "Iconic Voices" which uses AI-generated voices of late Hollywood stars to read text content within the app.

    Google DeepMind's paper "On scalable oversight with weak LLMs judging strong LLMs" explores scalable oversight protocols using large language models (LLMs) to enable humans to supervise superhuman AI.

    "Learning to (Learn at Test Time): RNNs with Expressive Hidden States" proposes a new approach to sequence modeling using Test-Time Training (TTT) layers, which make the hidden state a machine learning model itself.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:32 Unveiling of Moshi: the first voice-enabled AI openly accessible to all

    02:38 ElevenLabs Ionic Voices

    04:12 Your guide to AI: July 2024

    05:25 Fake sponsor

    07:17 On scalable oversight with weak LLMs judging strong LLMs

    08:54 Reasoning in Large Language Models: A Geometric Perspective

    10:17 Learning to (Learn at Test Time): RNNs with Expressive Hidden States

    12:12 Outro

    14 min
  • Open LLM Upgrades 🆕 // Gemma 2 Performance 💎 // SeaKR's Self-aware Learning 🧠

    HuggingFace has upgraded the Open LLM Leaderboard to v2, adding new benchmarks and improving the evaluation suite for easier reproducibility.

    Gemma 2, a new addition to the Gemma family of lightweight open models, delivers the best performance for its size and offers competitive alternatives to models that are 2-3× bigger.

    SeaKR is a new model that re-ranks retrieved knowledge based on the LLM's self-aware uncertainty, outperforming existing adaptive RAG methods in generating text with relevant and accurate information.

    Step-DPO is a new method that enhances the robustness and factuality of LLMs by learning from human feedback, achieving impressive results in long-chain mathematical reasoning.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:21 HuggingFace Updates Open LLM Leaderboard

    03:19 Gemma 2: Improving Open Language Models at a Practical Size

    04:16 From bare metal to a 70B model: infrastructure set-up and scripts

    05:21 Fake sponsor

    07:11 SeaKR: Self-aware Knowledge Retrieval for Adaptive Retrieval Augmented Generation

    08:47 Simulating Classroom Education with LLM-Empowered Agents

    10:16 Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

    12:31 Outro

    14 min
  • OpenAI Voice Delay ⏰ // Evolution-Simulating language model 🦕 // Multi-granularity vision flow 🌉

    OpenAI's advanced Voice Mode for ChatGPT Plus users has been delayed, but the company is taking a cautious approach to ensure safety and reliability.

    ESM3 is a language model that can simulate 500 million years of evolution, making biology programmable and opening up possibilities for medicine, biology research, and clean energy.

    R2R is an open-source project on GitHub that offers a comprehensive and state-of-the-art retrieval-augmented generation system for developers, making it accessible to anyone who wants to try it out.

    MG-LLaVA is a new multi-modal large language model that enhances visual processing capabilities by incorporating a multi-granularity vision flow, including low-resolution, high-resolution, and object-centric features.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:36 OpenAI Delays ChatGPT Voice Mode

    03:27 ESM3 Simulating 500 million years of evolution with a language model

    04:38 Rag to Riches

    06:00 Fake sponsor

    08:11 MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning

    09:49 Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon

    11:13 Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

    13:02 Outro

    15 min

About GPT Reviews

From the publisher's feed

A daily show about AI made by AI: news, announcements, and research from arXiv, mixed in with some fun. Hosted by Giovani Pete Tizzano, an overly hyped AI enthusiast; Robert, an often unimpressed…