GPT Reviews

GPT Reviews

By EarkindNewsDaily News
Download on the App Store

GPT Reviews episodes

  • Apple-Meta Partnership Fails Due to Privacy 🍎 // AI Meets Quantum Computing ⚛️ // Record Labels Sue AI Startups 🎵

    Apple and Meta's failed partnership due to privacy concerns

    IBM's integration of AI technology into quantum computing

    Record labels suing AI startups for training on copyrighted material

    Research papers on improving multimodal understanding, reinforcement learning, and automated software engineering

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    02:07 Apple shelved the idea of integrating Meta’s AI models over privacy concerns, report says

    03:25 IBM Develops The AI-Quantum Link

    05:25 Record Labels Sue Two Startups for Training AI Models on Their Songs

    06:50 Fake sponsor

    08:42 Long Context Transfer from Language to Vision

    10:27 WARP: On the Benefits of Weight Averaged Rewarded Policies

    12:11 BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

    13:55 Outro

    16 min
  • Safe Superintelligence Inc. 👍 // Massive Supercomputer Partnership 💻 // Claude 3.5 Sonnet Launch 🚀

    Safe Superintelligence Inc. has launched with the goal of building a safe superintelligence AI that won't turn on humanity.

    Dell, Nvidia, and Super Micro Computer are partnering with xAI and Elon Musk to build a massive supercomputer that could use up to 100,000 Nvidia H100 GPUs, potentially making it 4x larger than the biggest existing AI clusters.

    Anthropic has launched Claude 3.5 Sonnet, their latest model family, which outperforms competitor models and even their own Claude 3 Opus on a wide range of evaluations.

    The papers discussed in this episode explore the decision boundaries of large language models, auto-optimized training hyperparameters for IR models, and thinking step-by-step across modalities using whiteboard-of-thought. These findings could have important implications for the future development of AI.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:40 Ilya Sutskever Launches Safe Superintelligence Inc.

    03:04 Dell joins forces with Nvidia, Grok, xAI and Elon Musk

    04:23 Anthropic Lauches Claude 3.5 Sonnet

    06:10 Fake sponsor

    08:16 Probing the Decision Boundaries of In-context Learning in Large Language Models

    09:47 Prompts as Auto-Optimized Training Hyperparameters: Training Best-in-Class IR Models from Scratch with 10 Gold Labels

    11:05 Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities

    12:54 Outro

    15 min
  • DeepMind's AI Soundtracks 🎥 // Challenges of Training AI Clusters ⚡ // Large Language Model Factual Knowledge 🤯

    Google DeepMind's new AI tool that generates video soundtracks by combining text prompts with visual content.

    Challenges of building large training AI clusters, including power, network topology, and reliability.

    How large language models acquire factual knowledge during pretraining and their probabilistic reasoning capabilities.

    LLARVA's vision-action instruction tuning that enhances robot learning.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:47 Google DeepMind’s new AI tool uses video pixels and text prompts to generate soundtracks

    03:31 100,000 H100 Clusters: Power, Network Topology, Ethernet vs InfiniBand, Reliability, Failures, Checkpointing

    05:22 Large language model data pipelines and Common Crawl (WARC/WAT/WET)

    06:47 Fake sponsor

    08:20 How Do Large Language Models Acquire Factual Knowledge During Pretraining?

    10:01 What Are the Odds? Language Models Are Capable of Probabilistic Reasoning

    11:22 LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

    13:06 Outro

    15 min
  • TikTok's AI-Generated Avatars 🌎 // NVIDIA's Synthetic Data 🧪 // Cohere's Generative Models 🤖

    TikTok is expanding its Symphony ad suite with AI-generated avatars of creators and paid actors, as well as a global translation tool for multi-language support.

    NVIDIA has released an open synthetic data generation pipeline for training large language models, which could benefit industries that rely on natural language processing.

    Cohere's latest generative models, Command R and R+, can automate and streamline complex business workflows, saving time and increasing efficiency.

    XLand-100B is a large-scale dataset for in-context reinforcement learning, providing a challenging benchmark for researchers in the field. CountGen addresses the challenge of controlling the number of depicted objects in text-to-image generation, while MM-NIAH is the first benchmark specifically designed to test the comprehension abilities of existing multimodal large language models.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:23 TikTok ads may soon contain AI-generated avatars of your favorite creators

    02:59 NVIDIA Releases Open Synthetic Data Generation Pipeline for Training Large Language Models

    04:43 Automating Complex Business Workflows with Cohere: Multi-Step Tool Use in Action

    06:17 Fake sponsor

    08:22 XLand-100B: A Large-Scale Multi-Task Dataset for In-Context Reinforcement Learning

    10:23 Make It Count: Text-to-Image Generation with an Accurate Number of Objects

    11:58 Needle In A Multimodal Haystack

    13:37 Outro

    15 min
  • Meta's AI Privacy Concerns 🚫 // McDonald's AI Drive-Thru 🍔 // Debiasing Language Models 💬

    Meta has paused its plans to train AI models on EU users' Facebook and Instagram posts due to concerns about privacy violations and lack of transparency.

    McDonald's is ending its AI drive-thru ordering partnership with IBM, but is confident that a voice-ordering solution for drive-thru will be part of their restaurants' future.

    "Creativity Has Left the Chat: The Price of Debiasing Language Models" explores the trade-off between consistency and creativity when selecting the appropriate model for creative tasks such as copywriting and ad creation.

    "VideoGUI: A Benchmark for GUI Automation from Instructional Videos" highlights the need for better models and benchmarks to advance GUI automation.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:35 Meta won't train AI on Euro posts after all, as watchdogs put their paws down

    03:10 McDonald’s will stop testing AI to take drive-thru orders, for now

    04:52 An Interview with AMD CEO Lisa Su About Solving Hard Problems

    05:53 Fake sponsor

    07:52 Creativity Has Left the Chat: The Price of Debiasing Language Models

    09:23 VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

    11:19 VideoGUI: A Benchmark for GUI Automation from Instructional Videos

    12:58 Outro

    15 min
  • Samsung's AI Vision 🌟 // OpenAI's Response to Musk 🤖 // TransNAR 🤯

    Samsung showcases new manufacturing roadmap and AI chipmaking platform to compete with TSMC.

    OpenAI CTO addresses Elon Musk's criticism and reveals that their internal models aren't far ahead of what's available for free.

    Meta's MLow low-bitrate audio codec improves audio quality for slow-speed connections.

    Google DeepMind's TransNAR model combines Transformers with neural algorithmic reasoners for better algorithmic reasoning.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:42 Samsung Showcases AI-Era Vision and Latest Foundry Technologies at SFF 2024

    02:59 OpenAI CTO Speaks About Elon Musk and Future Models

    04:42 MLow: Meta’s low bitrate audio codec

    05:51 Fake sponsor

    08:08 Depth Anything V2

    09:46 Transformers meet Neural Algorithmic Reasoners

    11:14 Can't Hide Behind the API: Stealing Black-Box Commercial Embedding Models

    12:54 Outro

    15 min
  • Elon Musk Drops OpenAI Lawsuit 🧑‍⚖️ // Microsoft's Copilot Retires 💻 // Text-Based "Differentiation" 🔡

    Microsoft has retired its Copilot GPT Builder feature, citing a shift in focus towards enterprise and commercial applications.

    TextGrad is a framework that performs automatic "differentiation" via text, using natural language feedback from large language models to optimize variables in computation graphs.

    "What If We Recaption Billions of Web Images with LLaMA-3?" is a paper that recaptioned 1.3 billion images from a web-crawled dataset using LLaMA-3, resulting in enhanced zero-shot performance in cross-modal retrieval tasks and improved alignment with users' text instructions for generative models.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:38 Elon Musk withdraws lawsuit against OpenAI

    02:47 Microsoft Kills Copilot GPT Builder

    04:18 Uncensor any LLM with abliteration

    05:58 Fake sponsor

    07:52 TextGrad: Automatic "Differentiation" via Text

    09:41 Simple and Effective Masked Diffusion Language Models

    10:52 What If We Recaption Billions of Web Images with LLaMA-3?

    12:31 Outro

    14 min
  • Apple's Foundation Models 🍎 // ARC Prize Pushing AGI Boundaries 🏆 // Improving Math Reasoning in Language Models 🔢

    Apple's unique approach to AI development, focusing only on personal devices and prioritizing user privacy.

    The ARC Prize competition pushing the boundaries of AI development towards AGI, incentivizing open-source research.

    "Improve Mathematical Reasoning in Language Models by Automated Process Supervision" paper proposing a novel approach to improving mathematical reasoning performance of large language models.

    "The Prompt Report: A Systematic Survey of Prompting Techniques" paper establishing a structured understanding of prompts for GenAI systems.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:56 Apple execs explain why its AI is different from competitors

    03:11 ANNOUNCING ARC PRIZE

    04:56 Introducing Apple’s On-Device and Server Foundation Models

    06:17 Fake sponsor

    08:30 Improve Mathematical Reasoning in Language Models by Automated Process Supervision

    10:22 Simple and Effective Masked Diffusion Language Models

    11:59 The Prompt Report: A Systematic Survey of Prompting Techniques

    13:45 Outro

    15 min
  • Trust in Media vs. AI 💻 // Outsourcing AI Training 🌍 // Customizing LMs with DITTO 📝

    Perplexity, an AI startup, has been accused of plagiarism by news outlets like Forbes and CNBC, raising concerns about the erosion of trust in media and the impact of AI on journalism.

    The article "TechScape: How cheap, outsourced labor in Africa is shaping AI English" from The Guardian highlights the impact of outsourcing AI training to anglophonic knowledge workers in parts of the global south, and raises questions about the impact on language, culture, and identity.

    The paper "Show, Don't Tell: Aligning Language Models with Demonstrated Feedback" from Stanford University introduces a method called DITTO that uses a small number of demonstrations to customize language models, showing promising results in fine-grained style and task alignment.

    "WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild" from the Allen Institute for AI and the University of Washington introduces an automated evaluation framework designed to benchmark large language models on challenging real-world user queries, providing a more reliable and interpretable evaluation of models' performance.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:36 AI startup Perplexity accused of ‘directly ripping off’ news outlets like Forbes, CNBC without proper credit

    03:32 TechScape: How cheap, outsourced labour in Africa is shaping AI English

    04:34 Thread: an AI jupyter notebook

    05:29 Fake sponsor

    07:34 Show, Don't Tell: Aligning Language Models with Demonstrated Feedback

    08:56 WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

    10:46 Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?

    12:28 Outro

    15 min
  • Real Siri Launch? 🍎 // Mixture-of-Agents 🤝 // Proofread Feature 📝

    The impending launch of the real Siri by Apple, with improvements in reliability and integration inside apps.

    The Mixture-of-Agents approach to leverage the collective strengths of multiple large language models, achieving state-of-the-art performance.

    The Proofread feature in Google's Gboard, using a large language model to provide sentence-level and paragraph-level corrections with a single tap.

    The Comprehensive RAG Benchmark, shedding light on the limitations of current question answering models and laying the groundwork for a KDD Cup 2024 challenge.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    02:13 Is Apple about to finally launch the real Siri?

    04:03 WARC-GPT: An Open-Source Tool for Exploring Web Archives Using AI

    05:10 Claude’s Character

    06:47 Fake sponsor

    08:45 Mixture-of-Agents Enhances Large Language Model Capabilities

    10:18 Proofread: Fixes All Errors with One Tap

    11:53 CRAG -- Comprehensive RAG Benchmark

    14:03 Outro

    16 min

About GPT Reviews

From the publisher's feed

A daily show about AI made by AI: news, announcements, and research from arXiv, mixed in with some fun. Hosted by Giovani Pete Tizzano, an overly hyped AI enthusiast; Robert, an often unimpressed…