GPT Reviews

GPT Reviews

By EarkindNewsDaily News
Download on the App Store

GPT Reviews episodes

  • Apple's Smarter Siri 🍎 // Long-context models 🧐 // Natural Language Uncertainty 👀

    Apple is making strides in AI with their own model called Ajax and improvements to Siri, including making large language models faster and more efficient.

    "In-Context Learning with Long-Context Models: An In-Depth Exploration" explores a training method for long-context models called in-context learning and its effectiveness.

    "WildChat: 1M ChatGPT Interaction Logs in the Wild" offers a diverse dataset of user-chatbot interactions for researchers to study and fine-tune instruction-following models.

    "I'm Not Sure, But...": Examining the Impact of Large Language Models' Uncertainty Expression on User Reliance and Trust" investigates the impact of large language models on user reliance and trust, and the potential harm of overreliance. The study found that using natural language expressions of uncertainty can reduce overreliance on LLMs.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:34 Better Siri is coming: what Apple’s research says about its AI plans

    03:19 Your guide to AI: May 2024

    04:18 How LLMs Work, Explained Without Math

    05:37 Fake sponsor

    07:06 In-Context Learning with Long-Context Models: An In-Depth Exploration

    08:42 WildChat: 1M ChatGPT Interaction Logs in the Wild

    10:36 "I'm Not Sure, But...": Examining the Impact of Large Language Models' Uncertainty Expression on User Reliance and Trust

    12:55 Outro

    15 min
  • AI Spokeswoman for Ukraine 🗣️ // Anthropic iOS App 📱 // Bootstrapping Language Model Agents 🚀

    Ukraine introduces an AI-generated digital spokesperson for their Ministry of Foreign Affairs, named 'Victoriya Shi', who will deliver pre-prepared official statements on behalf of the ministry.

    Anthropic releases a mobile app version of their Claude AI models, including a new paid plan called Claude Team for group usage.

    BAGEL is a new method for bootstrapping language model agents without human supervision, which quickly converts the initial distribution of trajectories towards those that are well-described by natural language.

    CodeIt is a self-improvement method for language models that helps them improve their performance on complex reasoning tasks, achieving state-of-the-art performance and outperforming existing neural and symbolic baselines.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:25 Ukraine Unveils AI-generated Foreign Ministry Spokeswoman

    03:01 Anthropic finally releases a Claude mobile app

    04:51 Apple's Tiny LLMs, Amazon Rethinks Cashier-Free Stores, Predicting Scientific Discoveries

    06:46 Fake sponsor

    08:17 BAGEL: Bootstrapping Agents by Guiding Exploration with Language

    09:41 A Careful Examination of Large Language Model Performance on Grade School Arithmetic

    11:18 CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay

    13:09 Outro

    15 min
  • Amazon's Q 🤖 // Microsoft's OpenAI Investment 💰 // Global AI Math Championship 🏆

    Amazon has launched Q, an AI-powered assistant for businesses and developers that offers advanced capabilities such as code generation, testing, debugging, reasoning, and agents for step-by-step planning. 

    Microsoft's $1 billion investment in OpenAI was triggered by fears of falling behind Google in the AI race. The investment has helped Microsoft catch up and be seen as more of a leader in AI, with OpenAI's models integrated into their products. 

    A new dataset for the Global Artificial Intelligence Championship Math 2024 has been created, consisting of 387 math problems curated by professional math problem writers from prestigious institutions. 

    Three AI research papers were discussed, including a new approach to evaluating large language models using a panel of diverse models, a method for real-time, controllable motion generation, and the use of ranked list truncation for large language model-based re-ranking.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:36 Amazon Q, a generative AI-powered assistant for businesses and developers

    03:08 Microsoft’s OpenAI investment was triggered by Google fears, emails reveal

    05:12 A Dataset for The Global Artificial Intelligence Championship Math 2024

    06:21 Fake sponsor

    08:25 Ranked List Truncation for Large Language Model-based Re-Ranking

    10:04 Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

    11:35 MotionLCM: Real-time Controllable Motion Generation via Latent Consistency Model

    13:16 Outro

    15 min
  • Cohere on Amazon 🚀 // Big Tech Lobbying Frenzy 💼 // Multi-Token Prediction & KANs 🤖

    Cohere Command R & R+ now available on Amazon for enterprise-grade workloads and multilingual support.

    Big tech companies dominating AI lobbying efforts in Washington, potentially leading to weak regulations.

    Multi-token prediction proposed as a new way of training large language models, resulting in higher sample efficiency and faster inference.

    KANs, a new type of neural network with learnable activation functions on edges or weights, outperform MLPs in accuracy and interpretability, and can help scientists discover mathematical and physical laws.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:54 Cohere Command R & R+ now available on Amazon

    03:25 There’s an AI Lobbying Frenzy in Washington. Big Tech Is Dominating

    05:22 THE 150X PGVECTOR SPEEDUP: A YEAR-IN-REVIEW

    06:31 Fake sponsor

    08:04 Better & Faster Large Language Models via Multi-token Prediction

    09:55 KAN: Kolmogorov-Arnold Networks

    11:51 Iterative Reasoning Preference Optimization

    13:43 Outro

    15 min
  • OpenAI X Financial Times 📰 // GitHub Copilot Workspace 💻 // Memary: Long-term Memory for Agents 🧠

    OpenAI partners with the Financial Times to enhance ChatGPT with their award-winning journalism and develop new AI products and features for FT readers.

    GitHub announces the technical preview of GitHub Copilot Workspace, a Copilot-native developer environment that could revolutionize the way developers work.

    Memary, an open-source long-term memory system for autonomous agents, solves the problem of limited context windows for agents by allowing them to store a large corpus of information in knowledge graphs and retrieve only relevant information for meaningful responses.

    The papers discussed in this episode showcase the latest advancements in AI research, including AdvPrompter, HaLo-NeRF, and PLLaVA, which address issues related to large language models, digital exploration of large-scale tourist landmarks, and video understanding.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:51 We’re bringing the Financial Times’ world-class journalism to ChatGPT

    02:54 GitHub Copilot Workspace: Welcome to the Copilot-native developer environment

    04:43 memary: Open-Source Longterm Memory for Autonomous Agents

    05:55 Fake sponsor

    07:44 AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs

    09:26 HaLo-NeRF: Learning Geometry-Guided Semantics for Exploring Unconstrained Photo Collections

    11:17 PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

    13:18 Outro

    15 min
  • Chinese AI Model Beats GPT-4 🇨🇳 // OpenAI on iOS 18 🍎 // Data-Efficient LLMs 🤖

    SenseTime's new AI model, SenseNova 5.0, beats GPT-4 Turbo across key benchmarks, suggesting China's AI may be closer to competing with the US than previously thought.

    Apple is in talks with OpenAI to potentially integrate their features into iOS 18, which could trigger a new era of AI adoption.

    "Toward Inference-optimal Mixture-of-Expert Large Language Models" proposes a new scaling law for MoE-based LLMs to efficiently scale without sacrificing performance.

    "How to Train Data-Efficient LLMs" investigates data-efficient approaches for pre-training LLMs, which can significantly reduce the amount of data needed to train LLMs.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:30 Chinese AI model bests GPT-4 Turbo

    02:35 Apple Intensifies Talks With OpenAI for iPhone Generative AI Features

    04:17 OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

    05:33 Fake sponsor

    07:55 Toward Inference-optimal Mixture-of-Expert Large Language Models

    09:21 Scaling Laws For Dense Retrieval

    11:01 How to Train Data-Efficient LLMs

    12:50 Outro

    15 min
  • Investments Pay Off for MSFT 💰 // Apple's Language Models 🍎 // Improved Language Search 🔎

    Microsoft's investment in AI is paying off, with a 17% jump in revenue and a 20% increase in profit for the first three months of the year.

    Apple has released eight small AI language models aimed at on-device use, using a "layer-wise scaling strategy" to improve performance and transparency.

    Multi-Head Mixture-of-Experts is a new approach to address issues with Sparse Mixtures of Experts, outperforming existing models on three different tasks.

    Stream of Search (SoS) is a new technique for teaching language models to search, resulting in improved search accuracy and the ability to solve previously unsolved problems.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:28 Microsoft Reports Rising Revenues as A.I. Investments Bear Fruit

    03:14 Apple releases eight small AI language models aimed at on-device use

    05:00 Fake sponsor

    07:01 Multi-Head Mixture-of-Experts

    08:43 Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Perfect Reasoners

    10:30 Stream of Search (SoS): Learning to Search in Language

    12:31 Outro

    14 min
  • Meta's Stock Plunge 💸 // TSMC's A16 Process 🚀 // Instruction Hierarchy Boosting LLMs 📈

    Meta's aggressive AI investments have caused a 13% plunge in their stock, threatening to wipe out almost $163 billion from their market value.

    TSMC's new A16 manufacturing process promises to outperform its predecessor, N2P, by a significant margin, with an up to 10% higher clock rate at the same voltage and a 15% - 20% lower power consumption at the same frequency and complexity.

    The Instruction Hierarchy proposes a data generation method to demonstrate hierarchical instruction following behavior, which drastically increases robustness for LLMs against attacks.

    SPLATE is a lightweight adaptation of the ColBERTv2 model that improves the efficiency of late interaction retrieval, particularly for running ColBERT on CPU environments.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:27 Meta’s stock plunges on ‘aggressive’ AI spending plans

    02:49 TSMC unveils 1.6nm process technology with backside power delivery, rivals Intel's competing design

    04:48 tiny-gpu

    05:59 Fake sponsor

    07:35 The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

    08:43 A Reproducibility Study of PLAID

    10:18 SPLATE: Sparse Late Interaction Retrieval

    12:00 Outro

    14 min
  • Perplexity's Funding 🦄 // NVIDIA acquires Run:ai 🏎️ // Llama-3 on LM Leaderboard 🧐

    Perplexity becomes an AI unicorn with a new $63 million funding round.

    NVIDIA acquires Run:ai, an Israeli startup that provides Kubernetes-based workload management and orchestration software for AI computing resources.

    Llama-3 language model reaches the top-5 of the LM arena leaderboard.

    New AI research papers explore efficient language models, LLMs that can read your minds, and mixtures of experts. 

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:44 Perplexity becomes an AI unicorn with new $63 million funding round

    03:20 NVIDIA to Acquire GPU Orchestration Software Provider Run:ai

    05:24 Llama 3 on top-5 of LM arena leaderboard

    06:48 Fake sponsor

    08:54 OpenELM: An Efficient Language Model Family with Open-source Training and Inference Framework

    10:32 SnapKV: LLM Knows What You are Looking for Before Generation

    12:37 Multi-Head Mixture-of-Experts

    14:37 Outro

    16 min
  • Phi-3 from Microsoft 💻 // SoftBank Invests $1B in Nvidia 🤑 // HuggingFace's FineWeb Dataset 🌐

    Microsoft has launched its smallest AI model yet, the Phi-3 Mini, which is designed to be smaller and cheaper to run than its larger counterparts.

    SoftBank plans to invest nearly $1 billion in Nvidia's chips to bolster its computing facilities and develop its own generative AI, giving Japan a strong domestic player in the AI space.

    HuggingFace has released FineWeb, a dataset consisting of more than 15 trillion tokens of cleaned and deduplicated English web data from CommonCrawl, which outperforms models trained on other commonly used high-quality web datasets.

    The papers discussed in this episode cover topics such as extending embedding models for long context retrieval, automating graphic design using large multimodal models, and Microsoft's innovative approach to training the Phi-3 Mini AI model.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:35 Microsoft launches Phi-3, its smallest AI model yet

    03:10 SoftBank will reportedly invest nearly $1 billion in AI push, tapping Nvidia’s chips

    05:11 HuggingFace Releases FineWeb: 15 Trillion tokens to train on

    06:02 Fake sponsor

    08:15 Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

    09:42 LongEmbed: Extending Embedding Models for Long Context Retrieval

    11:04 Graphic Design with Large Multimodal Model

    12:53 Outro

    15 min

About GPT Reviews

From the publisher's feed

A daily show about AI made by AI: news, announcements, and research from arXiv, mixed in with some fun. Hosted by Giovani Pete Tizzano, an overly hyped AI enthusiast; Robert, an often unimpressed…