GPT Reviews

GPT Reviews

By EarkindNewsDaily News
Download on the App Store

GPT Reviews episodes

  • Nvidia's Stock Struggles 📉 // Meta's AI Hallucinations 🤖 // Superconducting Microprocessors ⚡

    This episode dives into Nvidia's stock struggles amid rising competition, while also unpacking Meta's AI blunders and the implications of "hallucinations" in tech. We explore cutting-edge superconducting microprocessors that promise unprecedented energy efficiency and highlight groundbreaking AI research, including eavesdropping techniques and advancements in reinforcement learning.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:50 Nvidia Sank Again Today -- Time to Buy the Artificial Intelligence (AI) Growth Stock Hand Over Fist?

    03:09 Meta blames hallucinations after its AI said Trump rally shooting didn’t happen

    04:52 Superconducting Microprocessors? Turns Out They're Ultra-Efficient

    06:07 Fake sponsor

    07:48 Deep-TEMPEST: Using Deep Learning to Eavesdrop on HDMI from its Unintended Electromagnetic Emanations

    09:22 SAPG: Split and Aggregate Policy Gradients

    10:45 MindSearch: Mimicking Human Minds Elicits Deep AI Searcher

    12:44 Outro

    15 min
  • Google's Gemma 2 vs. GPT-3.5 ⚔️ // Black Forest Labs' Flux Model 🌲 // Ethical Concerns in AI 🚨

    This episode dives into Google’s Gemma 2, which claims to outperform GPT-3.5 while tackling responsible AI practices. We explore Black Forest Labs' Flux model, featuring 12 billion parameters and tailored versions for various users. Olivia sheds light on the ethical concerns surrounding the resurgence of pseudoscience in machine learning, particularly physiognomy. Lastly, Belinda reviews critical research on AI safety, advocating for clearer metrics to prevent misleading claims about safety advancements.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:37 Google’s tiny AI model bests GPT-3.5

    02:48 Announcing Flux by Black Forest Labs: The Next Leap in Text-to-Image Models

    04:28 The reanimation of pseudoscience in machine learning and its ethical repercussions

    06:06 Fake sponsor

    08:04 MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts

    09:55 Spectra: A Comprehensive Study of Ternary, Quantized, and FP16 Language Models

    11:41 Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

    13:33 Outro

    15 min
  • Apple's AI Feature Delay 📅 // SAM 2 Object Segmentation 🖼️ // Google's TPU Chips Shift ⚡
    Apple’s delay in releasing AI features until October could affect iPhone 16 sales and customer excitement. The tech giant’s choice to use Google’s TPU chips instead of Nvidia marks a significant shift in AI hardware competition. Meta’s SAM 2 introduces groundbreaking real-time object segmentation with zero-shot generalization, revolutionizing visual content interaction. Additionally, Sony AI’s research presents a cost-effective approach to training diffusion models, democratizing access to advanced AI technology.
    Timestamps:
    00:34 Introduction
    01:54 Apple Intelligence Won't Be Released Until October
    03:09 Apple used Google's chips to train two AI models, research paper shows
    04:44 A Visual Guide to Quantization
    05:38 Introducing SAM 2: The next generation of Meta Segment Anything Model for videos and images
    06:41 Fake sponsor
    08:46 Stretching Each Dollar: Diffusion Training from Scratch on a Micro-Budget
    10:28 Theia: Distilling Diverse Vision Foundation Models for Robot Learning
    12:27 Outro
    15 min
  • OpenAI's SearchGPT 🧐 // AI in Math Olympiad 🏅 // Unreliable AI Existential Risk 🔍

    OpenAI's new prototype, SearchGPT, promises to combine AI smarts with real-time web information to make search easier.

    AI has achieved silver-medal standards at the International Mathematical Olympiad, raising questions about the future of mathematics and the role of AI in solving complex problems.

    The reliability of AI existential risk probabilities is called into question in a thought-provoking article, challenging the authority we often assign to these forecasts and calling for more scrutiny.

    Three fascinating papers from UNC Chapel Hill, Google DeepMind, and a collaboration between Caltech and NVIDIA explore advancements in theorem proving, balancing fast and slow planning, and aligning large language models with Best-of-N distillation. These papers could transform the way we approach complex problems with language models and streamline the development of LLMs.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:54 OpenAI Announces SearchGPT

    03:15 AI achieves silver-medal standard solving International Mathematical Olympiad problems

    04:55 AI existential risk probabilities are too unreliable to inform policy

    06:25 Fake sponsor

    08:21 LeanDojo: Theorem Proving with Retrieval-Augmented Language Models

    10:10 System-1.x: Learning to Balance Fast and Slow Planning with Language Models

    12:01 BOND: Aligning LLMs with Best-of-N Distillation

    13:43 Outro

    16 min
  • Mistral Large 2 🌍 // Memphis Supercluster 💻 // Emergence in Complex Systems 🧩

    Mistral Large 2 release with advanced features and multilingual support.

    Elon Musk's announcement of the Memphis Supercluster for creating the world's most powerful AI.

    Discussion of emergence in complex systems and the MINT-1T dataset for training large multimodal models.

    Introduction of OpenDevin, an open platform for developing AI agents and MOMAland, a benchmark framework for multi-objective multi-agent reinforcement learning.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:39 Mistral Large 2 Release

    03:01 Elon Musk Announces Memphis Supercomputer

    04:48 The Puzzle of How Large-Scale Order Emerges in Complex Systems

    06:22 Fake sponsor

    08:37 MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens

    10:16 OpenDevin: An Open Platform for AI Software Developers as Generalist Agents

    11:53 MOMAland: A Set of Benchmarks for Multi-Objective Multi-Agent Reinforcement Learning

    13:31 Outro

    15 min
  • Llama 3.1 Unveiled 🦙 // Alphabet's 14% Revenue Growth 📈 // MovieDreamer Revolutionizes Video 🎬

    This episode features the introduction of Llama 3.1, Meta's cutting-edge AI model with remarkable flexibility and extensive language support. We delve into Alphabet's impressive 14% revenue growth, highlighting the increasing demand for AI infrastructure in cloud computing. The System-1.x Planner is explored, demonstrating its innovative balance between fast and slow planning modes, leading to enhanced performance. Finally, we discuss MovieDreamer, a groundbreaking model that elevates video generation by ensuring narrative coherence and high visual quality in long-form content.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:39 Introducing Llama 3.1: Our most capable models to date

    02:59 Alphabet revenue jump shows no sign of AI denting search business

    04:36 Open Source AI Is the Path Forward

    05:40 Fake sponsor

    07:41 System-1.x: Learning to Balance Fast and Slow Planning with Language Models

    09:31 KAN or MLP: A Fairer Comparison

    11:08 MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequence

    12:53 Outro

    15 min
  • Meta's Llama 3.1 vs. GPT-4o 🤯 // OpenAI's own AI chips 🧐 // SlowFast-LLaVA for Video LLMs 🎬

    Meta's upcoming Llama 3.1 models could outperform the current state-of-the-art closed-source LLM model, OpenAI's GPT-4o.

    OpenAI is planning to develop its own AI chip to optimize performance and potentially supercharge their progress towards AGI.

    Apple's SlowFast-LLaVA is a new training-free video large language model that captures both detailed spatial semantics and long-range temporal context in video without exceeding the token budget of commonly used LLMs.

    Google's Conditioned Language Policy (CLP) framework is a general framework that builds on techniques from multi-task training and parameter-efficient finetuning to develop steerable models that can trade-off multiple conflicting objectives at inference time.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:28 LLAMA 405B Performance Leaked

    03:01 OpenAI Wants Its Own AI Chips

    04:25 Towards more cooperative AI safety strategies

    06:01 Fake sponsor

    07:35 SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

    09:17 AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?

    10:56 Conditioned Language Policy: A General Framework for Steerable Multi-Objective Finetuning

    12:46 Outro

    15 min
  • Claude for Android 🤖 // AI for Material Sciences ⚡ // TinkerBird Disrupts RAG Workflows 🐦

    Claude for Android is now available, bringing AI-powered assistance to a wider audience.

    MIT researchers have developed a new machine-learning framework that can predict materials' thermal properties up to 1,000 times faster than other AI-based techniques, potentially improving energy efficiency.

    TinkerBird, a vector database designed for efficient storage and retrieval of high-dimensional vectors, is disrupting traditional RAG workflows and eliminating roundtrip delays associated with client-server models.

    ChatQA 2, a Llama3-based model from NVIDIA, bridges the gap between open-access LLMs and leading proprietary models in long-context understanding and retrieval-augmented generation capabilities, while Stable Audio Open, an open-weights text-to-audio model from Stability AI, showcases potential for high-quality stereo sound synthesis at 44.1kHz.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:34 Claude for Android is here

    02:50 AI method radically speeds predictions of materials’ thermal properties

    04:44 TinkerBird

    06:10 Fake sponsor

    08:10 ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities

    09:54 Stable Audio Open

    11:28 Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders

    13:54 Outro

    16 min
  • OpenAI's GPT-4o mini 💰 // NVIDIA's Mistral NeMo 12B 🚀 // Transcribro speech recognition 🎤

    OpenAI has released their newest model, GPT-4o mini, which is more cost-efficient and excels in mathematical reasoning and coding tasks.

    NVIDIA's Mistral NeMo 12B is a state-of-the-art language model with unprecedented accuracy and enterprise-grade support.

    A new speech recognition keyboard and service for Android called Transcribro has been developed, which is private and on-device.

    Research papers explore the impact of vocabulary size on language model scaling, the use of large datastores for retrieval-based language models, and a method for generating long sequences of views of a cityscape using AI and computer vision.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:40 OpenAI Announces GPT 4o mini

    03:11 Mistral AI and NVIDIA Unveil Mistral NeMo 12B, a Cutting-Edge Enterprise AI Model

    05:28 Transcribro: Private and on-device speech recognition keyboard and service for Android

    06:43 Fake sponsor

    08:49 Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies

    10:19 Scaling Retrieval-Based Language Models with a Trillion-Token Datastore

    11:49 Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffusion

    13:26 Outro

    15 min
  • Copyright Infringement in AI Training 🚫 // Open-Source AI Models 🤖 // NVIDIA's Open-Source Transition 🆕

    Apple, Nvidia, Anthropic, and Salesforce caught using content without creators' consent for AI training.

    Mistral AI launches two new open-source models, Codestral Mamba and Mathstral, with impressive capabilities.

    NVIDIA transitions to fully open-source GPU kernel modules, offering new capabilities and easy switching for users.

    Exciting research papers include Ref-AVS for multimodal object segmentation, Qwen2-Audio for large-scale audio-language modeling, and DiT-MoE for scalable language modeling and image generation.

    Contact:  [email protected]

    Timestamps:

    00:34 Introduction

    01:27 Apple, Nvidia, Anthropic Used Thousands of Swiped YouTube Videos to Train AI

    02:46 Mistral's New Open Source Models

    04:09 NVIDIA Transitions Fully Towards Open-Source GPU Kernel Modules

    05:37 Fake sponsor

    07:15 Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes

    08:47 Qwen2-Audio Technical Report

    10:49 Scaling Diffusion Transformers to 16 Billion Parameters

    12:21 Outro

    14 min

About GPT Reviews

From the publisher's feed

A daily show about AI made by AI: news, announcements, and research from arXiv, mixed in with some fun. Hosted by Giovani Pete Tizzano, an overly hyped AI enthusiast; Robert, an often unimpressed…