GPT Reviews

GPT Reviews

By EarkindNewsDaily News
Download on the App Store

GPT Reviews episodes

  • AI for Musicians 🎶 // GPT-5 Upgrade 📈 // RewardBench Evaluation 🧑‍💻

    Soundry AI promises to be a game-changer for musicians with its superior flexibility and versatility compared to standard sample libraries.

    OpenAI is set to release GPT-5, an improved version of the AI language model that powers ChatGPT, which could represent a notable advancement for OpenAI.

    RewardBench, a benchmark dataset and codebase for evaluating reward models, provides a standardized way to evaluate reward models on a range of tasks, including chat, reasoning, and safety.

    DepthFM's Fast Monocular Depth Estimation with Flow Matching is a promising direction for the field of monocular depth estimation, with its generative approach and state-of-the-art performance on standard benchmarks.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:33 Soundry, AI for Musicians, by Musicians

    02:55 GPT-5 might arrive this summer as a “materially better” update to ChatGPT

    04:55 The Reddits

    06:38 Fake sponsor

    08:33 RewardBench: Evaluating Reward Models for Language Modeling

    10:11 Evaluating Frontier Models for Dangerous Capabilities

    11:25 DepthFM: Fast Monocular Depth Estimation with Flow Matching

    12:57 Outro

    14 min
  • Microsoft hires Suleyman 💻 // NVIDIA's GR00T humanoid robot 🤖 // uBlockOrigin AI blocklist 🚫

    Microsoft hires Mustafa Suleyman to lead AI division, including Copilot, an AI tool that helps programmers write better code.

    NVIDIA announces Project GR00T, a multimodal AI system that enables advanced humanoid robots to learn skills and interact with the real world, with partnerships from leading robotics companies.

    LLMLingua-2 and Agent-FLAN are two cutting-edge AI research papers that address the challenges of prompt compression and incorporating agent ability into large language models, respectively.

    uBlockOrigin and uBlacklist's huge AI blocklist is a useful resource for those who want to eradicate AI-generated content from their search results.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:52 Microsoft hires Suleyman to lead AI

    03:08 NVIDIA Announces Project GR00T Foundation Model for Humanoid Robots and Major Isaac Robotics Platform Update

    04:45 uBlockOrigin & uBlacklist Huge AI Blocklist

    05:55 Fake sponsor

    07:50 LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression

    09:22 mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

    11:14 Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models

    13:17 Outro

    15 min
  • Sam Altman on GPT-5 💰 // Compute as Currency 💻 // Vid2Robot 🤖

    Saudi Arabia's $40 billion investment in AI, making them the largest investor in the field.

    OpenAI's CEO Sam Altman's prediction that compute will be the currency of the future.

    Vid2Robot, a video-based learning framework that directly produces robot actions given a video demonstration of a manipulation task and the robot's current visual observations.

    TexDreamer, the first zero-shot multimodal high-fidelity 3D human texture generation model.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:50 Saudi Arabia Plans $40 Billion Push Into Artificial Intelligence

    03:14 Sam Altman on GPT-5 and the future of AI

    04:46 Price Per Part in Lego Sets

    05:57 Fake sponsor

    07:52 Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

    09:35 mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

    11:31 TexDreamer: Towards Zero-Shot High-Fidelity 3D Human Texture Generation

    13:20 Outro

    15 min
  • Nvidia's Blackwell B200 🚀 // Stable Video 3D 💥 // Apple's AI Partnership Talks 🍎

    Nvidia's Blackwell B200 GPU is the world's most powerful chip for AI, offering up to 20 petaflops of FP4 horsepower and reducing cost and energy consumption.

    Stable Video 3D is a generative model that can be used for commercial purposes and improves the quality of 3D meshes generated directly from novel views.

    Apple is in talks with potential partners like Google to integrate their AI models into the iPhone and the broader OS ecosystem, which could be a major win for both companies.

    The latest AI research includes papers on the security of API-protected LLMs, optimizing attention modules for processing long sequences, and super-resolution of anime images.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:40 Nvidia reveals Blackwell B200 GPU, the ‘world’s most powerful chip’ for AI

    03:46 Introducing Stable Video 3D: Quality Novel View Synthesis and 3D Generation from Single Images

    05:39 Apple's AI plans reportedly could involve a partner like Google

    06:51 Fake sponsor

    08:52 Logits of API-Protected LLMs Leak Proprietary Information

    10:35 BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences

    12:19 APISR: Anime Production Inspired Real-World Anime Super-Resolution

    13:59 Outro

    16 min
  • Figure AI Robots 🤖 // OpenAI Leaks GPT-5 🤫 // Scaling Language Models 📈

    Figure, a leading AI robotics company, is making significant advancements in creating robots that can perceive their environment, make decisions, and take action, all in a way that aligns with human expectations.

    OpenAI may have accidentally leaked details about a new AI model called GPT-4.5 Turbo, which could level the playing field with Google's AI model Gemini.

    Two papers explore the development and evaluation of large language models (LLMs) for code-related tasks, and propose simple and scalable strategies to continually pre-train LLMs to save on compute.

    Another paper investigates scaling in the over-trained regime and relates language model perplexity to downstream task performance via a power law, providing useful insights into how language models can be scaled and evaluated more effectively.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:28 Figure AI

    03:02 Did OpenAI just accidentally leak the next big ChatGPT upgrade?

    04:48 Gradio's Grog

    05:47 Fake sponsor

    07:37 Simple and Scalable Strategies to Continually Pre-train Large Language Models

    09:31 LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

    11:11 Language models scale reliably with over-training and on downstream tasks

    13:07 Outro

    15 min
  • Grok-1 Released 🤖 // Landmark EU AI Act 🇪🇺 // Multimodal LLM Importance 🔍

    X AI has released their 314 billion parameter Mixture-of-Experts model, Grok-1, which is currently the largest language model that has been publicly released and could be used for a variety of tasks.

    The European Parliament has passed a landmark AI act that bans certain AI applications and requires strict obligations for high-risk AI, positioning itself as the global standard for regulation.

    The paper "MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training" explores the importance of Multimodal Large Language Models (MLLM) and how a careful mix of image-caption, interleaved image-text, and text-only data is crucial for achieving state-of-the-art few-shot results across multiple benchmarks.

    The paper "Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews" examines the impact of large language models on scientific peer review and suggests that LLM-generated text can affect the quality and fairness of peer review.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:24 Open Release of Grok-1

    02:43 Claude 3 Haiku

    04:34 EU passes landmark AI act

    06:15 The Rest of the World Disappears’: Claire Voisin on Mathematical Creativity

    07:59 Fake sponsor

    09:54 MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

    11:30 Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews

    13:19 Outro

    15 min
  • Devin: First AI Software Engineer 🤖 // Google Gemini AI on Elections 🚫 // WorkArena Benchmark 💼

    Introducing Devin, the first AI software engineer that can plan and execute complex engineering tasks requiring thousands of decisions.

    Google's AI chatbot won't answer questions about upcoming elections to prevent inaccurate or misleading responses.

    WorkArena, a benchmark measuring the ability of large language model-based agents to perform tasks that align with the daily work of knowledge workers using enterprise software systems.

    Synth$^2$, a novel approach that leverages Large Language Models (LLMs) and image generation models to create synthetic image-text pairs for efficient and effective Visual-Language Model (VLM) training.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:59 Introducing Devin, the first AI software engineer

    03:48 Google won’t let you use its Gemini AI to answer questions about an upcoming election in your country

    05:35 AI Datacenter Energy Dilemma - Race for AI Datacenter Space

    06:45 Fake sponsor

    09:08 Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM

    10:26 WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

    11:52 Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings

    13:38 Outro

    15 min
  • Language Model Advancements 🔍 // Meta's AI Investment 💰 // Model Security Risks 🔒

    The latest research in AI language models, including algorithmic progress and multistep consistency models.

    A new LLM model called Command-R, designed for large-scale production workloads.

    The announcement of Meta's investment in AI infrastructure, including two 24k GPU clusters and plans for continued growth.

    A discussion on the potential security risks of model-stealing attacks on language models.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:42 Command-R: Retrieval Augmented Generation at Production Scale

    03:48 Building Meta’s GenAI Infrastructure

    05:41 MLX Server

    06:37 Fake sponsor

    08:45 Algorithmic progress in language models

    10:23 Multistep Consistency Models

    11:35 Stealing Part of a Production Language Model

    13:31 Outro

    15 min
  • Nvidia's Dominance 🏆 // AI & Cryptocurrency's Energy Demands 💡 // OpenAI's Transformer Debugger 🔍

    Nvidia's CEO claims that even free AI chips from competitors can't beat Nvidia's GPUs, highlighting the company's dominance in the AI industry.

    The energy demands of AI and cryptocurrency are discussed, with concerns raised about the potential consequences of their growing electricity needs.

    OpenAI's Superalignment team has developed a powerful new tool called Transformer Debugger, which allows researchers to investigate specific behaviors of small language models.

    Three AI research papers are discussed, including a model-stealing attack on black-box production language models, a study questioning the effectiveness of cosine-similarity in measuring semantic similarity, and the latest version of SPLADE, a library for ranking natural language queries in information retrieval systems. 

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:49 Jensen Huang says even free AI chips from his competitors can't beat Nvidia's GPUs

    03:42 The Obscene Energy Demands of A.I.

    05:03 Transformer Debugger

    06:27 Fake sponsor

    08:07 Stealing Part of a Production Language Model

    09:37 Is Cosine-Similarity of Embeddings Really About Similarity?

    11:05 SPLADE-v3: New baselines for SPLADE

    12:46 Outro

    14 min
  • Apple's AI Release Plans 🍎 // Sam Altman Returns to Board 🤝 // Novel Transformer-Based Architectures 🧐

    Apple's upcoming AI releases in 2024, including a new Siri upgrade and a rumored "AppleGPT" LLM.

    Sam Altman's return to OpenAI's board of directors after an investigation into his ouster.

    Research on high-resolution image synthesis using rectified flow transformers and a novel transformer-based architecture for text-to-image generation.

    A new metric called FENICE for evaluating factual inconsistencies in automatically generated text summaries.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:45 Apple Preparing AI Releases for 2024

    03:03 Sam Altman rejoins OpenAI board of directors as investigation into his ouster comes to a close

    05:11 Your guide to AI: March 2024

    06:45 Fake sponsor

    08:41 Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

    10:50 FENICE: Factuality Evaluation of summarization based on Natural language Inference and Claim Extraction

    12:19 Poly-View Contrastive Learning

    14:16 Outro

    16 min

About GPT Reviews

From the publisher's feed

A daily show about AI made by AI: news, announcements, and research from arXiv, mixed in with some fun. Hosted by Giovani Pete Tizzano, an overly hyped AI enthusiast; Robert, an often unimpressed…