ThursdAI - The top AI news from the past week

ThursdAI - The top AI news from the past week

By From Weights & Biases, Join AI Evangelist Alex Volkov and a panel of experts to cover everything important that happened in the world of AI from the past weekNewsTechnologyTech News
Download on the App Store

ThursdAI - The top AI news from the past week episodes

  • ๐Ÿ“† ThursdAI - Nov 7 - Video version, full o1 was given and taken away, Anthropic price hike-u, halloween ๐Ÿ’€ recap & more AI news

    ๐Ÿ‘‹ Hey all, this is Alex, coming to you from the very Sunny California, as I'm in SF again, while there is a complete snow storm back home in Denver (brrr).

    I flew here for the Hackathon I kept telling you about, and it was glorious, we had over 400 registered, over 200 approved hackers, 21 teams submitted incredible projects ๐Ÿ‘ You can follow some of these here

    I then decided to stick around and record the show from SF, and finally pulled the plug and asked for some budget, and I present, the first ThursdAI, recorded from the newly minted W&B Podcast studio at our office in SF ๐ŸŽ‰

    This isn't the only first, today also, for the first time, all of the regular co-hosts of ThursdAI, met on video for the first time, after over a year of hanging out weekly, we've finally made the switch to video, and you know what? Given how good AI podcasts are getting, we may have to stick around with this video thing! We played one such clip from a new model called hertz-dev, which is a <10B model for full duplex audio.

    Given that today's episode is a video podcast, I would love for you to see it, so here's the timestamps for the chapters, which will be followed by the TL;DR and show notes in raw format. I would love to hear from folks who read the longer form style newsletters, do you miss them? Should I bring them back? Please leave me a comment ๐Ÿ™ (I may send you a survey)

    This was a generally slow week (for AI!! not for... ehrm other stuff) and it was a fun podcast! Leave me a comment about what you think about this new format.

    Chapter Timestamps

    00:00 Introduction and Agenda Overview

    00:15 Open Source LLMs: Small Models

    01:25 Open Source LLMs: Large Models

    02:22 Big Companies and LLM Announcements

    04:47 Hackathon Recap and Community Highlights

    18:46 Technical Deep Dive: HertzDev and FishSpeech

    33:11 Human in the Loop: AI Agents

    36:24 Augmented Reality Lab Assistant

    36:53 Hackathon Highlights and Community Vibes

    37:17 Chef Puppet and Meta Ray Bans Raffle

    37:46 Introducing Fester the Skeleton

    38:37 Fester's Performance and Community Reactions

    39:35 Technical Insights and Project Details

    42:42 Big Companies API Updates

    43:17 Haiku 3.5: Performance and Pricing

    43:44 Comparing Haiku and Sonnet Models

    51:32 XAI Grok: New Features and Pricing

    57:23 OpenAI's O1 Model: Leaks and Expectations

    01:08:42 Transformer ASIC: The Future of AI Hardware

    01:13:18 The Future of Training and Inference Chips

    01:13:52 Oasis Demo and Etched AI Controversy

    01:14:37 Nisten's Skepticism on Etched AI

    01:19:15 Human Layer Introduction with Dex

    01:19:24 Building and Managing AI Agents

    01:20:54 Challenges and Innovations in AI Agent Development

    01:21:28 Human Layer's Vision and Future

    01:36:34 Recap and Closing Remarks

    ThursdAI - Recaps of the most high signal AI weekly spaces is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

    Show Notes and Links:

    * Interview

    * Dexter Horthy (X) from HumanLayer

    * Open Source LLMs

    * SmolLM2: the new, best, and open 1B-parameter language mode (X)

    * Meta released MobileLLM (125M, 350M, 600M, 1B) (HF)

    * Tencent Hunyuan Large - 389B X 52B (Active) MoE (X, HF, Paper)

    * Big CO LLMs + APIs

    * OpenAI buys and opens chat.com

    * Anthropic releases Claude Haiku 3.5 via API (X, Blog)

    * OpenAI drops o1 full - and pulls it back (but not before it got Jailbroken)

    * X.ai now offers $25/mo free of Grok API credits (X, Platform)

    * Etched announces Sonu - first Transformer ASIC - 500K tok/s (etched)

    * PPXL is not valued at 9B lol

    * This weeks Buzz

    * Recap of SF Hackathon w/ AI Tinkerers (X)

    * Fester the Halloween Toy aka Project Halloweave videos from trick or treating (X, Writeup)

    * Voice & Audio

    * Hertz-dev - 8.5B conversation audio gen (X, Blog )

    * Fish Agent v0.1 3B - Speech to Speech model (HF, Demo)

    * AI Art & Diffusion & 3D

    * FLUX 1.1 [pro] is how HD - 4x resolution (X, blog)

    Full Transcription for convenience below:



    This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
    1 hr 39 min
  • ๐Ÿ“† ThursdAI - Spooky Halloween edition with Video!

    Hey everyone, Happy Halloween! Alex here, coming to you live from my mad scientist lair! For the first ever, live video stream of ThursdAI, I dressed up as a mad scientist and had my co-host, Fester the AI powered Skeleton join me (as well as my usual cohosts haha) in a very energetic and hopefully entertaining video stream!

    Since it's Halloween today, Fester (and I) have a very busy schedule, so no super length ThursdAI news-letter today, as we're still not in the realm of Gemini being able to write a decent draft that takes everything we talked about and cover all the breaking news, I'm afraid I will have to wish you a Happy Halloween and ask that you watch/listen to the episode.

    The TL;DR and show links from today, don't cover all the breaking news but the major things we saw today (and caught live on the show as Breaking News) were, ChatGPT now has search, Gemini has grounded search as well (seems like the release something before Google announces it streak from OpenAI continues).

    Here's a quick trailer of the major things that happened:

    This weeks buzz - Halloween AI toy with Weave

    In this weeks buzz, my long awaited Halloween project is finally live and operational!

    I've posted a public Weave dashboard here and the code (that you can run on your mac!) here

    Really looking forward to see all the amazing costumers the kiddos come up with and how Gemini will be able to respond to them, follow along!

    ThursdAI - Recaps of the most high signal AI weekly spaces is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

    Ok and finally my raw TL;DR notes and links for this week. Happy halloween everyone, I'm running off to spook the kiddos (and of course record and post about it!)

    ThursdAI - Oct 31 - TL;DR

    TL;DR of all topics covered:

    * Open Source LLMs:

    * Microsoft's OmniParser: SOTA UI parsing (MIT Licensed) ๐•

    * Groundbreaking model for web automation (MIT license).

    * State-of-the-art UI parsing and understanding.

    * Outperforms GPT-4V in parsing web UI.

    * Designed for web automation tasks.

    * Can be integrated into various development workflows.

    * ZhipuAI's GLM-4-Voice: End-to-end Chinese/English speech ๐•

    * End-to-end voice model for Chinese and English speech.

    * Open-sourced and readily available.

    * Focuses on direct speech understanding and generation.

    * Potential applications in various speech-related tasks.

    * Meta releases LongVU: Video LM for long videos ๐•

    * Handles long videos with impressive performance.

    * Uses DINOv2 for downsampling, eliminating redundant scenes.

    * Fuses features using DINOv2 and SigLIP.

    * Select tokens are passed to Qwen2/Llama-3.2-3B.

    * Demo and model are available on HuggingFace.

    * Potential for significant advancements in video understanding.

    * OpenAI new factuality benchmark (Blog, Github)

    * Introducing SimpleQA: new factuality benchmark

    * Goal: high correctness, diversity, challenging for frontier models

    * Question Curation: AI trainers, verified by second trainer

    * Quality Assurance: 3% inherent error rate

    * Topic Diversity: wide range of topics

    * Grading Methodology: "correct", "incorrect", "not attempted"

    * Model Comparison: smaller models answer fewer correctly

    * Calibration Measurement: larger models more calibrated

    * Limitations: only for short, fact-seeking queries

    * Conclusion: drive research on trustworthy AI

    * Big CO LLMs + APIs:

    * ChatGPT now has Search! (X)

    * Grounded search results in browsing the web

    * Still hallucinates

    * Reincarnation of Search GPT inside ChatGPT

    * Apple Intelligence Launch: Image features for iOS 18.2 [๐•]( Link not provided in source material)

    * Officially launched for developers in iOS 18.2.

    * Includes Image Playground and Gen Moji.

    * Aims to enhance image creation and manipulation on iPhones.

    * GitHub Universe AI News: Co-pilot expands, new Spark tool ๐•

    * GitHub Co-pilot now supports Claude, Gemini, and OpenAI models.

    * GitHub Spark: Create micro-apps using natural language.

    * Expanding the capabilities of AI-powered coding tools.

    * Copilot now supports multi-file edits in VS Code, similar to Cursor, and faster code reviews.

    * GitHub Copilot extensions are planned for release in 2025.

    * Grok Vision: Image understanding now in Grok ๐•

    * Finally has vision capabilities (currently via ๐•, API coming soon).

    * Can now understand and explain images, even jokes.

    * Early version, with rapid improvements expected.

    * OpenAI advanced voice mode updates (X)

    * 70% cheaper in input tokens because of automatic caching (X)

    * Advanced voice mode is now on desktop app

    * Claude this morning - new mac / pc App

    * This week's Buzz:

    * My AI Halloween toy skeleton is greeting kids right now (and is reporting to Weave dashboard)

    * Vision & Video:

    * Meta's LongVU: Video LM for long videos ๐• (see Open Source LLMs for details)

    * Grok Vision on ๐•: ๐• (see Big CO LLMs + APIs for details)

    * Voice & Audio:

    * MaskGCT: New SoTA Text-to-Speech ๐•

    * New open-source state-of-the-art text-to-speech model.

    * Zero-shot voice cloning, emotional TTS, long-form synthesis, variable speed synthesis, bilingual (Chinese & English).

    * Available on Hugging Face.

    * ZhipuAI's GLM-4-Voice: End-to-end Chinese/English speech ๐• (see Open Source LLMs for details)

    * Advanced Voice Mode on Desktops: ๐• (See Big CO LLMs + APIs for details).

    * AI Art & Diffusion: (See Red Panda in "This week's Buzz" above)

    * Redcraft Red Panda: new SOTA image diffusion ๐•

    * High-performing image diffusion model, beating Black Forest Labs Flux.

    * 72% win rate, higher ELO than competitors.

    * Creates SVG files, editable as vector files.

    * From Redcraft V3.

    * Tools:

    * Bolt.new by StackBlitz: In-browser full-stack dev environment ๐•

    * Platform for prompting, editing, running, and deploying full-stack apps directly in your browser.

    * Uses WebContainers.

    * Supports npm, Vite, Next.js, and integrations with Netlify, Cloudflare, and SuperBase.

    * Free to use.

    * Jina AI's Meta-Prompt: Improved LLM Codegen ๐•



    This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
    1 hr 50 min
  • ๐Ÿ“… ThursdAI - Oct 24 - Claude 3.5 controls your PC?! Talking AIs with ๐Ÿฆพ, Multimodal Weave, Video Models mania + more AI news from this ๐Ÿ”ฅ week.

    Hey all, Alex here, coming to you from the (surprisingly) sunny Seattle, with just a mind-boggling week of releases. Really, just on Tuesday there was so much news already! I had to post a recap thread, something I do usually after I finish ThursdAI!

    From Anthropic reclaiming close-second sometimes-first AI lab position + giving Claude the wheel in the form of computer use powers, to more than 3 AI video generation updates with open source ones, to Apple updating Apple Intelligence beta, it's honestly been very hard to keep up, and again, this is literally part of my job!

    But once again I'm glad that we were able to cover this in ~2hrs, including multiple interviews with returning co-hosts ( Simon Willison came back, Killian came back) so definitely if you're only a reader at this point, listen to the show!

    Ok as always (recently) the TL;DR and show notes at the bottom (I'm trying to get you to scroll through ha, is it working?) so grab a bucket of popcorn, let's dive in ๐Ÿ‘‡

    ThursdAI - Recaps of the most high signal AI weekly spaces is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

    Claude's Big Week: Computer Control, Code Wizardry, and the Mysterious Case of the Missing Opus

    Anthropic dominated the headlines this week with a flurry of updates and announcements. Let's start with the new Claude Sonnet 3.5 (really, they didn't update the version number, it's still 3.5 tho a different API model)

    Claude Sonnet 3.5: Coding Prodigy or Benchmark Buster?

    The new Sonnet model shows impressive results on coding benchmarks, surpassing even OpenAI's O1 preview on some. "It absolutely crushes coding benchmarks like Aider and Swe-bench verified," I exclaimed on the show. But a closer look reveals a more nuanced picture. Mixed results on other benchmarks indicate that Sonnet 3.5 might not be the universal champion some anticipated. My friend who has held back internal benchmarks was disappointed highlighting weaknesses in scientific reasoning and certain writing tasks. Some folks are seeing it being lazy-er for some full code completion, while the context window is now doubled from 4K to 8K! This goes to show again, that benchmarks don't tell the full story, so we wait for LMArena (formerly LMSys Arena) and the vibe checks from across the community.

    However it absolutely dominates in code tasks, that much is clear already. This is a screenshot of the new model on Aider code editing benchmark, a fairly reliable way to judge models code output, they also have a code refactoring benchmark

    Haiku 3.5 and the Vanishing Opus: Anthropic's Cryptic Clues

    Further adding to the intrigue, Anthropic announced Claude 3.5 Haiku! They usually provide immediate access, but Haiku remains elusive, saying that it's available by end of the month, which is very very soon. Making things even more curious, their highly anticipated Opus model has seemingly vanished from their website. "They've gone completely silent on 3.5 Opus," Simon Willison (๐•) noted, mentioning conspiracy theories that this new Sonnet might simply be a rebranded Opus? ๐Ÿ•ฏ๏ธ ๐Ÿ•ฏ๏ธ We'll make a summoning circle for new Opus and update you once it lands (maybe next year)

    Claude Takes Control (Sort Of): Computer Use API and the Dawn of AI Agents (๐•)

    The biggest bombshell this week? Anthropic's Computer Use. This isn't just about executing code; itโ€™s about Claude interacting with computers, clicking buttons, browsing the web, and yes, even ordering pizza! Killian Lukas (๐•), creator of Open Interpreter, returned to ThursdAI to discuss this groundbreaking development. "This stuff of computer useโ€ฆitโ€™s the same argument for having humanoid robots, the web is human shaped, and we need AIs to interact with computers and the web the way humans do" Killian explained, illuminating the potential for bridging the digital and physical worlds.

    Simon, though enthusiastic, provided a dose of realism: "It's incredibly impressiveโ€ฆbut also very much a V1, beta.โ€ Having tackled the setup myself, I agree; the current reliance on a local Docker container and virtual machine introduces some complexity and security considerations. However, seeing Claude fix its own Docker installation error was an unforgettably mindblowing experience. The future of AI agents is upon us, even if itโ€™s still a bit rough around the edges.

    Here's an easy guide to set it up yourself, takes 5 minutes, requires no coding skills and it's safely tucked away in a container.

    Big Tech's AI Moves: Apple Embraces ChatGPT, X.ai API (+Vision!?), and Cohere Multimodal Embeddings

    The rest of the AI world wasnโ€™t standing still. Apple made a surprising integration, while X.ai and Cohere pushed their platforms forward.

    Apple iOS 18.2 Beta: Siri Phones a Friend (ChatGPT)

    Apple, always cautious, surprisingly integrated ChatGPT directly into iOS. While Siri remainsโ€ฆwell, Siri, users can now effortlessly offload more demanding tasks to ChatGPT. "Siri is still stupid," I joked, "but can now ask it to write some stuff and it'll tell you, hey, do you want me to ask my much smarter friend ChatGPT about this task?" This approach acknowledges Siri's limitations while harnessing ChatGPTโ€™s power. The iOS 18.2 beta also includes GenMoji (custom emojis!) and Visual Intelligence (multimodal camera search) which are both welcome, tho I didn't really get the need of the Visual Intelligence (maybe I'm jaded with my Meta Raybans that already have this and are on my face most of the time) and I didn't get into the GenMoji waitlist still waiting to show you some custom emojis!

    X.ai API: Grok's Enterprise Ambitions and a Secret Vision Model

    Elon Musk's X.ai unveiled their API platform, focusing on enterprise applications with Grok 2 beta. They also teased an undisclosed vision model, and they had vision APIs for some folks who joined their hackathon. While these models are still not worth using necessarily, the next Grok-3 is promising to be a frontier model, and for some folks, it's relaxed approach to content moderation (what Elon is calling maximally seeking the truth) is going to be a convincing point for some!

    I just wish they added fun mode and access to real time data from X! Right now it's just the Grok-2 model, priced at a very non competative $15/mTok ๐Ÿ˜’

    Cohere Embed 3: Elevating Multimodal Embeddings (Blog)

    Cohere launched Embed 3, enabling embeddings for both text and visuals such as graphs and designs. "While not the first multimodal embeddings, when it comes from Cohere, you know it's done right," I commented.

    Open Source Power: JavaScript Transformers and SOTA Multilingual Models

    The open-source AI community continues to impress, making powerful models accessible to all.

    Massive kudos to Xenova (๐•) for the release of Transformers.js v3! The addition of WebGPU support results in a staggering "up to 100 times faster" performance boost for browser-based AI, dramatically simplifying local, private, and efficient model running. We also saw DeepSeekโ€™s Janus 1.3B, a multimodal image and text model, and Cohere For AI's Aya Expanse, supporting 23 languages.

    This Weekโ€™s Buzz: Hackathon Triumphs and Multimodal Weave

    On ThursdAI, we also like to share some of the exciting things happening behind the scenes.

    AI Chef Showdown: Second Place and Lessons Learned

    Happy to report that team Yes Chef clinched second place in a hackathon with an unconventional creation: a Gordon Ramsay-inspired robotic chef hand puppet, complete with a cloned voice and visual LLM integration. We bought and 3D printed and assembled an Open Source robotic arm, made it become a ventriloquist operator by letting it animate a hand puppet, and cloned Ramsey's voice. It was so so much fun to build, and the code is here

    Weave Goes Multimodal: Seeing and Hearing Your AI

    Even more exciting was the opportunity to leverage Weave's newly launched multimodal functionality. "Weave supports you to see and play back everything that's audio generated," I shared, emphasizing its usefulness in debugging our vocal AI chef.

    For a practical example, here's ALL the (NSFW) roasts that AI Chef has cooked me with, it's honestly horrifying haha. For full effect, turn on the background music first and then play the chef audio ๐Ÿ˜‚

    ๐Ÿ“ฝ๏ธ Video Generation Takes Center Stage: Mochi's Motion Magic and Runway's Acting Breakthrough

    Video models made a quantum leap this week, pushing the boundaries of generative AI.

    Genmo Mochi-1: Diffusion Transformers and Generative Motion

    Genmo's Ajay Jain (Genmo) joined ThursdAI to discuss Mochi-1, their powerful new diffusion transformer. "We really focused onโ€ฆprompt adherence and motion," he explained. Mochi-1's capacity to generate complex and realistic motion is truly remarkable, and with an HD version on its way, the future looks bright (and animated!). They also get bonus points for dropping a torrent link in the announcement tweet.

    So far this apache 2, 10B Diffusion Transformer is open source, but not for the GPU-poors, as it requires 4 GPUs to run, but apparently there was already an attempt to run in on one single 4090 which, Ajay highlighted was one of the reasons they open sourced it!

    Runway Act-One: AI-Powered Puppetry and the Future of Acting (blog)

    Ok this one absolutely seems bonkers! Runway unveiled Act-One! Forget just generating video from text; Act-One takes a driving video and character image to produce expressive and nuanced character performances. "It faithfully represents elements like eye-lines, micro expressions, pacing, and delivery," I noted, excited by the transformative potential for animation and filmmaking.

    So no need for rigging, for motion capture suites on faces of actors, Runway now, does this, so you can generate characters with Flux, and animate them with Act-One ๐Ÿ“ฝ๏ธ Just take a look at this insanity ๐Ÿ‘‡

    11labs Creative Voices: Prompting Your Way to the Perfect Voice

    11labs debuted an incredible feature: creating custom voices using only text prompts. Want a high-pitched squeak or a sophisticated British accent? Just ask. This feature makes bespoke voice creation significantly easier.

    I was really really impressed by this, as this is perfect for my Skeleton Halloween project! So far I struggled to get the voice "just right" between the awesome Cartesia voice that is not emotional enough, and the very awesome custom OpenAI voice that needs a prompt to act, and sometimes stops acting in the middle of a sentence.

    With this new Elevenlabs feature, I can describe the exact voice I want with a prompt, and then keep iterating until I find the perfect one, and then boom, it's available for me! Great for character creation, and even greater for the above Act-One model, as you can now generate a character with Flux, Drive the video with Act-one and revoice yourself with a custom prompted voice from 11labs! Which is exactly what I'm going to build for the next hackathon!

    If you'd like to support me in this journey, here's an 11labs affiliate link haha but I already got a yearly account so don't sweat it.

    AI Art & Diffusion Updates: Stable Diffusion 3.5, Ideogram Canvas, and OpenAI's Sampler Surprise

    The realm of AI art and diffusion models saw its share of action as well.

    Stable Diffusion 3.5 (Blog) and Ideogram Canvas: Iterative Improvements and Creative Control

    Stability AI launched Stable Diffusion 3.5, bringing incremental enhancements to image quality and prompt accuracy. Ideogram, meanwhile, introduced Canvas, a groundbreaking interface enabling mixing, matching, extending, and fine-tuning AI-generated artwork. This opens doors to unprecedented levels of control and creative expression.

    Midjourney also announced a web editor, and folks are freaking out, and I'm only left thinking, is MJ a bit a cult? There are so much offerings out there, but it seems like everything MJ releases gets tons more excitement from that part of X than other way more incredible stuff ๐Ÿค”

    Seattle Pic

    Ok wow that was a LOT of stuff to cover, honestly, the TL;DR for this week became so massive that I had to zoom out to take 1 screenshot of it all ,and I wasn't sure we'd be able to cover all of it!

    Massive massive week, super exciting releases, and the worst thing about this is, I barely have time to play with many of these!

    But I'm hoping to have some time during the Tinkerer AI hackathon we're hosting on Nov 2-3 in our SF office, limited spots left, so come and hang with me and some of the Tinkerers team, and maybe even win a Meta Rayban special Weave prize!

    RAW TL;DR + Show notes and links

    * Open Source LLMs

    * Xenova releases Transformers JS version 3 (X)

    * โšก WebGPU support (up to 100x faster than WASM)๐Ÿ”ข New quantization formats (dtypes)๐Ÿ› 120 supported architectures in total๐Ÿ“‚ 25 new example projects and templates๐Ÿค– Over 1200 pre-converted models๐ŸŒ Node.js (ESM + CJS), Deno, and Bun compatibility๐Ÿก A new home on GitHub and NPM

    * DeepSeek drops Janus 1.3B (X, HF, Paper)

    * DeepSeek releases Janus 1.3B ๐Ÿ”ฅ

    * ๐ŸŽจ Understands and generates both images and text

    * ๐Ÿ‘€Combines DeepSeek LLM 1.3B with SigLIP-L for vision

    * โœ‚๏ธ Decouples the vision encoding

    * Cohere for AI releases Aya expanse 8B, 32B (X, HF, Try it)

    * Aya Expanse is an open-weight research release of a model with highly advanced multilingual capabilities. It focuses on pairing a highly performant pre-trained Command family of models with the result of a yearโ€™s dedicated research from Cohere For AI, including data arbitrage, multilingual preference training, safety tuning, and model merging. The result is a powerful multilingual large language model serving 23 languages.

    * 23 languages

    * Big CO LLMs + APIs

    * New Claude Sonnet 3.5, Claude Haiku 3.5

    * New Claude absolutely crushes coding benchmarks like Aider and Swe-bench verified.

    * But I'm getting mixed signals from folks with internal benchmarks, as well as some other benches like Aidan Bench and Arc challenge in which it performs worse.

    * 8K output token limit vs 4K

    * Other folks swear by it, Skirano, Corbitt say it's an absolute killer coder

    * Haiku is 2x the price of 4o-mini and Flash

    * Anthropic Computer use API + docker (X)

    * Computer use is not new, see open interpreter etc

    * Adept has been promising this for a while, so was LAM from rabbit.

    * Now Anthropic has dropped a bomb on all these with a specific trained model to browse click and surf the web with a container

    * Examples of computer use are super cool, Corbitt built agent.exe which uses it to control your computer

    * Killian will join to talk about what this computer use means

    * Folks are trying to order food (like Anthropic shows in their demo of ordering pizzas for the team)

    * Claude launches code interpreter mode for claude.ai (X)

    * Cohere released Embed 3 for multimodal embeddings (Blog)

    * ๐Ÿ” Multimodal Embed 3: Powerful AI search model

    * ๐ŸŒ Unlocks value from image data for enterprises

    * ๐Ÿ” Enables fast retrieval of relevant info & assets

    * ๐Ÿ›’ Transforms e-commerce search with image search

    * ๐ŸŽจ Streamlines design process with visual search

    * ๐Ÿ“Š Improves data-driven decision making with visual insights

    * ๐Ÿ” Industry-leading accuracy and performance

    * ๐ŸŒ Multilingual support across 100+ languages

    * ๐Ÿค Partnerships with Azure AI and Amazon SageMaker

    * ๐Ÿš€ Available now for businesses and developers

    * X ai has a new API platform + secret vision feature (docs)

    * grok-2-beta $5.0 / $15.00 mtok

    * Apple releases IOS 18.2 beta with GenMoji, Visual Intelligence, ChatGPT integration & more

    * Siri is still stupid, but can now ask chatGPT to write s**t

    * This weeks Buzz

    * Got second place for the hackathon with our AI Chef that roasts you in the kitchen (X, Weave dash)

    * Weave is now multimodal and supports audio! (Weave)

    * Tinkerers Hackathon in less than a week!

    * Vision & Video

    * Genmo releases Mochi-1 txt2video model w/ Apache 2.0 license

    * Gen mo - generative motion

    * 10B DiT - diffusion transformer

    * 5.5 seconds video

    * Apache 2.0

    * Comparison thread between Genmo Mochi-1 and Hailuo

    * Genmo, the company behind Mochi 1, has raised $28.4M in Series A funding from various investors. Mochi 1 is an open-source video generation model that the company claims has "superior motion quality, prompt adherence and exceptional rendering of humans that begins to cross the uncanny valley." The company is open-sourcing their base 480p model, with an HD version coming soon.

    Summary Bullet Points:

    * Genmo announces $28.4M Series A funding

    * Mochi 1 is an open-source video generation model

    * Mochi 1 has "superior motion quality, prompt adherence and exceptional rendering of humans"

    * X is open-sourcing their base 480p Mochi 1 model

    * HD version of Mochi 1 is coming soon

    * Mochi 1 is available via Genmo's playground or as downloadable weights, or on Fal

    * Mochi 1 is licensed under Apache 2.0

    * Rhymes AI - Allegro video model (X)

    * Meta a bunch of releases - Sam 2.1, Spirit LM

    * Runway introduces puppetry video 2 video with emotion transfer (X)

    * The webpage introduces Act-One, a new technology from Runway that allows for the generation of expressive character performances using a single driving video and character image, without the need for motion capture or rigging. Act-One faithfully represents elements like eye-lines, micro expressions, pacing, and delivery in the final generated output. It can translate an actor's performance across different character designs and styles, opening up new avenues for creative expression.

    Summary in 10 Bullet Points:

    * Act-One is a new technology from Runway

    * It generates expressive character performances

    * Uses a single driving video and character image

    * No motion capture or rigging required

    * Faithfully represents eye-lines, micro expressions, pacing, and delivery

    * Translates performance across different character designs and styles

    * Allows for new creative expression possibilities

    * Works with simple cell phone video input

    * Replaces complex, multi-step animation workflows

    * Enables capturing the essence of an actor's performance

    * Haiper releases a new video model

    * Meta releases Sam 2.1

    * Key updates to SAM 2:

    * New data augmentation for similar and small objects

    * Improved occlusion handling

    * Longer frame sequences in training

    * Tweaks to positional encoding

    SAM 2 Developer Suite released:

    * Open source code package

    * Training code for fine-tuning

    * Web demo front-end and back-end code

    * Voice & Audio

    * OpenAI released custom voice support for chat completion API (X, Docs)

    * Pricing is still insane ($200/1mtok)

    * This is not just TTS, this is advanced voice mode!

    * The things you can ddo with them are very interesting, like asking for acting, or singing.

    * 11labs create voices with a prompt is super cool (X)

    * Meta Spirit LM: An open source language model for seamless speech and text integration (Blog, weights)

    * Meta Spirit LM is a multimodal language model that:

    * Combines text and speech processing

    * Uses word-level interleaving for cross-modality generation

    * Has two versions:

    * Base: uses phonetic tokens

    * Expressive: uses pitch and style tokens for tone

    * Enables more natural speech generation

    * Can learn tasks like ASR, TTS, and speech classification

    * MoonShine for audio

    * AI Art & Diffusion & 3D

    * Stable Diffusion 3.5 was released (X, Blog, HF)

    * including Stable Diffusion 3.5 Large and Stable Diffusion 3.5 Large Turbo.

    * table Diffusion 3.5 Medium will be released on October 29th.ย ย 

    * the permissive Stability AI Community License.ย 

    * ๐Ÿš€ Introducing Stable Diffusion 3.5 - powerful, customizable, and free models

    * ๐Ÿ” Improved prompt adherence and image quality compared to previous versions

    * โšก๏ธ Stable Diffusion 3.5 Large Turbo offers fast inference times

    * ๐Ÿ”ง Multiple variants for different hardware and use cases

    * ๐ŸŽจ Empowering creators to distribute and monetize their work

    * ๐ŸŒ Available for commercial and non-commercial use under permissive license

    * ๐Ÿ” Listening to community feedback to advance their mission

    * ๐Ÿ”„ Stable Diffusion 3.5 Medium to be released on October 29th

    * ๐Ÿค– Commitment to transforming visual media with accessible AI tools

    * ๐Ÿ”œ Excited to see what the community creates with Stable Diffusion 3.5

    * Ideogram released Canvas (X)

    * Canvas is a mix of Krea and Everart

    * Ideogram is a free AI tool for generating realistic images, posters, logos

    * Extend tool allows expanding images beyond original borders

    * Magic Fill tool enables editing specific image regions and details

    * Ideogram Canvas is a new interface for organizing, generating, editing images

    * Ideogram uses AI to enhance the creative process with speed and precision

    * Developers can integrate Ideogram's Magic Fill and Extend via the API

    * Privacy policy and other legal information available on the website

    * Ideogram is free-to-use, with paid plans offering additional features

    * Ideogram is available globally, with support for various browsers

    * OpenAI released a new sampler paper trying to beat diffusers (Blog)

    * Researchers at OpenAI have developed a new approach called sCM that simplifies the theoretical formulation of continuous-time consistency models, allowing them to stabilize and scale the training of these models for large datasets. The sCM approach achieves sample quality comparable to leading diffusion models, while using only two sampling steps - a 50x speedup over traditional diffusion models. Benchmarking shows sCM produces high-quality samples using less than 10% of the effective sampling compute required by other state-of-the-art generative models.The key innovation is that sCM models scale commensurately with the teacher diffusion models they are distilled from. As the diffusion models grow larger, the relative difference in sample quality between sCM and the teacher model diminishes. This allows sCM to leverage the advances in diffusion models to achieve impressive sample quality and generation speed, unlocking new possibilities for real-time, high-quality generative AI across domains like images, audio, and video.

    * ๐Ÿ” Simplifying continuous-time consistency models

    * ๐Ÿ”จ Stabilizing training for large datasets

    * ๐Ÿ” Scaling to 1.5 billion parameters on ImageNet

    * โšก 2-step sampling for 50x speedup vs. diffusion

    * ๐ŸŽจ Comparable sample quality to diffusion models

    * ๐Ÿ“Š Benchmarking against state-of-the-art models

    * ๐Ÿ—บ๏ธ Visualization of diffusion vs. consistency models

    * ๐Ÿ–ผ๏ธ Selected 2-step samples from 1.5B model

    * ๐Ÿ“ˆ Scaling sCM with teacher diffusion models

    * ๐Ÿ”ญ Limitations and future work

    * Midjourney announces an editor (X)

    * announces the release of two new features for Midjourney users - an image editor for uploaded images and

    * image re-texturing for exploring materials, surfacing, and lighting.

    * These features will initially be available only to yearly members, members who have been subscribers for the past 12 months, and members with at least 10,000 images.

    * The post emphasizes the need to give the community, human moderators, and AI moderation systems time to adjust to the new features

    * Tools

    PS : Subscribe to the newsletter and podcast, and I'll be back next week with more AI escapades! ๐Ÿซถ



    This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
    1 hr 57 min
  • ๐Ÿ“† ThursdAI - Oct 17 - Robots, Rockets, and Multi Modal Mania with open source voice cloning, OpenAI new voice API and more AI news

    Hey folks, Alex here from Weights & Biases, and this week has been absolutely bonkers. From robots walking among us to rockets landing on chopsticks (well, almost), the future is feeling palpably closer. And if real-world robots and reusable spaceship boosters weren't enough, the open-source AI community has been cooking, dropping new models and techniques faster than a Starship launch. So buckle up, grab your space helmet and noise-canceling headphones (weโ€™ll get to why those are important!), and let's blast off into this weekโ€™s AI adventures!

    TL;DR and show-notes + links at the end of the post ๐Ÿ‘‡

    Robots and Rockets: A Glimpse into the Future

    I gotta start with the real-world stuff because, let's be honest, it's mind-blowing. We had Robert Scoble (yes, the Robert Scoble) join us after attending the Tesla We, Robot AI event, reporting on Optimus robots strolling through crowds, serving drinks, and generally being ridiculously futuristic. Autonomous robo-taxis were also cruising around, giving us a taste of a driverless future.

    Robertโ€™s enthusiasm was infectious: "It was a vision of the future, and from that standpoint, it succeeded wonderfully." I couldn't agree more. While the market might have had a mini-meltdown (apparently investors aren't ready for robot butlers yet), the sheer audacity of Teslaโ€™s vision is exhilarating. These robots aren't just cool gadgets; they represent a fundamental shift in how we interact with technology and the world around us. And theyโ€™re learning fast. Just days after the event, Tesla released a video of Optimus operating autonomously, showcasing the rapid progress theyโ€™re making.

    And speaking of audacious visions, SpaceX decided to one-up everyone (including themselves) by launching Starship and catching the booster with Mechazilla โ€“ their giant robotic chopsticks (okay, technically a launch tower, but you get the picture). Waking up early with my daughter to watch this live was pure magic. As Ryan Carson put it, "It was magical watching thisโ€ฆ my kid who's 16โ€ฆ all of his friends are getting their imaginations lit by this experience." Thatโ€™s exactly what we need - more imagination and less doomerism! The future is coming whether we like it or not, and I, for one, am excited.

    Open Source LLMs and Tools: The Community Delivers (Again!)

    Okay, back to the virtual world (for now). This week's open-source scene was electric, with new model releases and tools that have everyone buzzing (and benchmarking like crazy!).

    * Nemotron 70B: Hype vs. Reality: NVIDIA dropped their Nemotron 70B instruct model, claiming impressive scores on certain benchmarks (Arena Hard, AlpacaEval), even suggesting it outperforms GPT-4 and Claude 3.5. As always, we take these claims with a grain of salt (remember Reflection?), and our resident expert, Nisten, was quick to run his own tests. The verdict? Nemotron is good, "a pretty good model to use," but maybe not the giant-killer some hyped it up to be. Still, kudos to NVIDIA for pushing the open-source boundaries. (Hugging Face, Harrison Kingsley evals)

    * Zamba 2 : Hybrid Vigor: Zyphra, in collaboration with NVIDIA, released Zamba 2, a hybrid Sparse Mixture of Experts (SME) model. We had Paolo Glorioso, a researcher from Ziphra, join us to break down this unique architecture, which combines the strengths of transformers and state space models (SSMs). He highlighted the memory and latency advantages of SSMs, especially for on-device applications. Definitely worth checking out if youโ€™re interested in transformer alternatives and efficient inference.

    * Zyda 2: Data is King (and Queen): Alongside Zamba 2, Zyphra also dropped Zyda 2, a massive 5 trillion token dataset, filtered, deduplicated, and ready for LLM training. This kind of open-source data release is a huge boon to the community, fueling the next generation of models. (X)

    * Ministral: Pocket-Sized Power: On the one-year anniversary of the iconic Mistral 7B release, Mistral announced two new smaller models โ€“ Ministral 3B and 8B. Designed for on-device inference, these models are impressive, but as always, Qwen looms large. While Mistral didnโ€™t include Qwen in their comparisons, early tests suggest Qwenโ€™s smaller models still hold their own. One point of contention: these Ministrals aren't as open-source as the original 7B, which is a bit of a bummer, with the 3B not being even released anywhere besides their platform. (Mistral Blog)

    * Entropix (aka Shrek Sampler): Thinking Outside the (Sample) Box: This one is intriguing! Entropix introduces a novel sampling technique aimed at boosting the reasoning capabilities of smaller LLMs. Nistenโ€™s yogurt analogy explains it best: itโ€™s about โ€œmarinatingโ€ the information and picking the best โ€œflavorโ€ (token) at the end. Early examples look promising, suggesting Entropix could help smaller models tackle problems that even trip up their larger counterparts. But, as with all shiny new AI toys, we're eagerly awaiting robust evals. Tim Kellog has an detailed breakdown of this method here

    * Gemma-APS: Fact-Finding Mission: Google released Gemma-APS, a set of models specifically designed for extracting claims and facts from text. While LLMs can already do this to some extent, a dedicated model for this task is definitely interesting, especially for applications requiring precise information retrieval. (HF)

    ๐Ÿ”ฅ OpenAI adds voice to their completion API (X, Docs)

    In the last second of the pod, OpenAI decided to grace us with Breaking News!

    Not only did they launch their Windows native app, but also added voice input and output to their completion APIs. This seems to be the same model as the advanced voice mode (and priced super expensively as well) and the one they used in RealTime API released a few weeks ago at DevDay.

    This is of course a bit slower than RealTime but is much simpler to use, and gives way more developers access to this incredible resource (I'm definitely planning to use this for ... things ๐Ÿ˜ˆ)

    This isn't their "TTS" or "STT (whisper) models, no, this is an actual omni model that understands audio natively and also outputs audio natively, allowing for things like "count to 10 super slow"

    I've played with it just now (and now it's after 6pm and I'm still writing this newsletter) and it's so so awesome, I expect it to be huge because the RealTime API is very curbersome and many people don't really need this complexity.

    This weeks Buzz - Weights & Biases updates

    Ok I wanted to send a completely different update, but what I will show you is, Weave, our observability framework is now also Multi Modal!

    This couples very well with the new update from OpenAI!

    So here's an example usage with today's announcement, I'm going to go through the OpenAI example and show you how to use it with streaming so you can get the audio faster, and show you the Weave multimodality as well ๐Ÿ‘‡

    You can find the code for this in this Gist and please give us feedback as this is brand new

    Non standard use-cases of AI corner

    This week I started noticing and collecting some incredible use-cases of Gemini and it's long context and multimodality and wanted to share with you guys, so we had some incredible conversations about non-standard use cases that are pushing the boundaries of what's possible with LLMs.

    Hrishi blew me away with his experiments using Gemini for transcription and diarization. Turns out, Gemini is not only great at transcription (it beats whisper!), itโ€™s also ridiculously cheaper than dedicated ASR models like Whisper, around 60x cheaper! He emphasized the unexplored potential of prompting multimodal models, adding, โ€œthe prompting on these thingsโ€ฆ is still poorly understood." So much room for innovation here!

    Simon Willison then stole the show with his mind-bending screen-scraping technique. He recorded a video of himself clicking through emails, fed it to Gemini Flash, and got perfect structured data in return. This trick isnโ€™t just clever; itโ€™s practically free, thanks to the ridiculously low cost of Gemini Flash. I even tried it myself, recording my X bookmarks and getting a near-perfect TLDR of the weekโ€™s AI news. The future of data extraction is here, and it involves screen recordings and very cheap (or free) LLMs.

    Here's Simon's example of how much this would cost him had he actually be charged for it. ๐Ÿคฏ

    Speaking of Simon Willison , he broke the news that NotebookLM has got an upgrade, with the ability to steer the speakers with custom commands, which Simon promptly used to ask the overview hosts to talk like Pelicans

    Voice Cloning, Adobe Magic, and the Quest for Real-Time Avatars

    Voice cloning also took center stage this week, with the release of F5-TTS. This open-source model performs zero-shot voice cloning with just a few seconds of audio, raising all sorts of ethical questions (and exciting possibilities!). I played a sample on the show, and it was surprisingly convincing (though not without it's problems) for a local model!

    This, combined with Hallo 2's (also released this week!) ability to animate talking avatars, has Wolfram Ravenwolf dreaming of real-time AI assistants with personalized faces and voices. The pieces are falling into place, folks.

    And for all you Adobe fans, Firefly Video has landed! This โ€œcommercially safeโ€ text-to-video and image-to-video model is seamlessly integrated into Premiere, offering incredible features like extending video clips with AI-generated frames. Photoshop also got some Firefly love, with mind-bending relighting capabilities that could make AI-generated images indistinguishable from real photographs.

    Wrapping Up:

    Phew, that was a marathon, not a sprint! From robots to rockets, open source to proprietary, and voice cloning to video editing, this week has been a wild ride through the ever-evolving landscape of AI. Thanks for joining me on this adventure, and as always, keep exploring, keep building, and keep pushing those AI boundaries. The future is coming, and itโ€™s going to be amazing.

    P.S. Donโ€™t forget to subscribe to the podcast and newsletter for more AI goodness, and if youโ€™re in Seattle next week, come say hi at the AI Tinkerers meetup. Iโ€™ll be demoing my Halloween AI toy โ€“ itโ€™s gonna be spooky!

    ThursdAI - Recaps of the most high signal AI weekly spaces is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

    TL;DR - Show Notes and Links

    * Open Source LLMs

    * Nvidia releases Llama 3.1-Nemotron-70B instruct: Outperforms GPT-40 and Anthropic Claude 3.5 on several benchmarks. Available on Hugging Face and Nvidia. (X, Harrison Eval)

    * Zamba2-7B: A hybrid Sparse Mixture of Experts model from Zyphra and Nvidia. Claims to outperform Mistral, Llama2, and Gemmas in the 58B weight class. (X, HF)

    * Zyda-2: 57B token dataset distilled from high-quality sources for training LLMs. Released by Zyphra and Nvidia. (X)

    * Ministral 3B & 8B - Mistral releases 2 new models for on device, claims SOTA (Blog)

    โ€ข Entropix aims to mimic advanced reasoning in small LLMs (Github, Breakdown)

    * Google releases Gemma-APS: A collection of Gemma models for text-to-propositions segmentation, distilled from Gemini Pro and fine-tuned on synthetic data. (HF)

    * Big CO LLMs + APIs

    * OpenAI ships advanced voice model in chat completions API endpoints with multimodality (X, Docs, My Example)

    * Amazon, Microsoft, Google all announce nuclear power for AI future

    * Yi-01.AI launches Yi-Lightning: A proprietary model accessible via API.

    * New Gemini API parameters: Google has shipped new Gemini API parameters, including logprobs, candidateCount, presencePenalty, seed, frequencyPenalty, and model_personality_in_response.

    * Google NotebookLM is no longer "experimental" and now allows for "steering" the hosts (Announcement)

    * XAI - GROK 2 and Grok2-mini are now available via API in OpenRouter - (X, OR)

    * This weeks Buzz (What I learned with WandB this week)

    * Weave is now MultiModal (supports audio and text!) (X, Github Example)

    * Vision & Video

    * Adobe Firefly Video: Adobe's first commercially safe text-to-video and image-to-video generation model. Supports prompt coherence. (X)

    * Voice & Audio

    * Ichigo-Llama3.1 Local Real-Time Voice AI: Improvements allow it to talk back, recognize when it can't comprehend input, and run on a single Nvidia 3090 GPU. (X)

    * F5-TTS: Performs zero-shot voice cloning with less than 15 seconds of audio, using audio clips to generate additional audio. (HF, Paper)

    * AI Art & Diffusion & 3D

    * RF-Inversion: Zero-shot inversion and editing framework for Flux, introduced by Litu Rout. Allows for image editing and personalization without training, optimization, or prompt-tuning. (X)

    * Tools

    * Fastdata: A library for synthesizing 1B tokens. (X)



    This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
    1 hr 36 min
  • ๐Ÿ“† ThursdAI - Oct 10 - Two Nobel Prizes in AI!? Meta Movie Gen (and sounds ) amazing, Pyramid Flow a 2B video model, 2 new VLMs & more AI news!

    Hey Folks, we are finally due for a "relaxing" week in AI, no more HUGE company announcements (if you don't consider Meta Movie Gen huge), no conferences or dev days, and some time for Open Source projects to shine. (while we all wait for Opus 3.5 to shake things up)

    This week was very multimodal on the show, we covered 2 new video models, one that's tiny and is open source, and one massive from Meta that is aiming for SORA's crown, and 2 new VLMs, one from our friends at REKA that understands videos and audio, while the other from Rhymes is apache 2 licensed and we had a chat with Kwindla Kramer about OpenAI RealTime API and it's shortcomings and voice AI's in general.

    ThursdAI - Recaps of the most high signal AI weekly spaces is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

    All right, let's TL;DR and show notes, and we'll start with the 2 Nobel prizes in AI ๐Ÿ‘‡

    * 2 AI nobel prizes

    * John Hopfield and Geoffrey Hinton have been awarded a Physics Nobel prize

    * Demis Hassabis, John Jumper & David Baker, have been awarded this year's #NobelPrize in Chemistry.

    * Open Source LLMs & VLMs

    * TxT360: a globally deduplicated dataset for LLM pre-training ( Blog, Dataset)

    * Rhymes Aria - 25.3B multimodal MoE model that can take image/video inputs Apache 2 (Blog, HF, Try It)

    * Maitrix and LLM360 launch a new decentralized arena (Leaderboard, Blog)

    * New Gradio 5 with server side rendering (X)

    * LLamaFile now comes with a chat interface and syntax highlighting (X)

    * Big CO LLMs + APIs

    * OpenAI releases MLEBench - new kaggle focused benchmarks for AI Agents (Paper, Github)

    * Inflection is still alive - going for enterprise lol (Blog)

    * new Reka Flash 21B - (X, Blog, Try It)

    * This weeks Buzz

    * We chatted about Cursor, it went viral, there are many tips

    * WandB releases HEMM - benchmarks of text-to-image generation models (X, Github, Leaderboard)

    * Vision & Video

    * Meta presents Movie Gen 30B - img and text to video models (blog, paper)

    * Pyramid Flow - open source img2video model MIT license (X, Blog, HF, Paper, Github)

    * Voice & Audio

    * Working with OpenAI RealTime Audio - Alex conversation with Kwindla from trydaily.com

    * Cartesia Sonic goes multilingual (X)

    * Voice hackathon in SF with 20K prizes (and a remote track) - sign up

    * Tools

    * LM Studio ships with MLX natively (X, Download)

    * UITHUB.com - turn any github repo into 1 long file for LLMs

    A Historic Week: TWO AI Nobel Prizes!

    This week wasn't just big; it was HISTORIC. As Yam put it, "two Nobel prizes for AI in a single week. It's historic." And he's absolutely spot on! Geoffrey Hinton, often called the "grandfather of modern AI," alongside John Hopfield, were awarded the Nobel Prize in Physics for their foundational work on neural networks - work that paved the way for everything we're seeing today. Think back propagation, Boltzmann machines โ€“ these are concepts that underpin much of modern deep learning. Itโ€™s about time they got the recognition they deserve!

    Yoshua Bengio posted about this in a very nice quote:

    @HopfieldJohn and @geoffreyhinton, along with collaborators, have created a beautiful and insightful bridge between physics and AI. They invented neural networks that were not only inspired by the brain, but also by central notions in physics such as energy, temperature, system dynamics, energy barriers, the role of randomness and noise, connecting the local properties, e.g., of atoms or neurons, to global ones like entropy and attractors. And they went beyond the physics to show how these ideas could give rise to memory, learning and generative models; concepts which are still at the forefront of modern AI research

    And Hinton's post-Nobel quote? Pure gold: โ€œIโ€™m particularly proud of the fact that one of my students fired Sam Altman." He went on to explain his concerns about OpenAI's apparent shift in focus from safety to profits. Spicy take! It sparked quite a conversation about the ethical implications of AI development and whoโ€™s responsible for ensuring its safe deployment. Itโ€™s a discussion we need to be having more and more as the technology evolves. Can you guess which one of his students it was?

    Then, not to be outdone, the AlphaFold team (Demis Hassabis, John Jumper, and David Baker) snagged the Nobel Prize in Chemistry for AlphaFold 2. This AI revolutionized protein folding, accelerating drug discovery and biomedical research in a way no one thought possible. These awards highlight the tangible, real-world applications of AI. It's not just theoretical anymore; it's transforming industries.

    Congratulations to all winners, and we gotta wonder, is this a start of a trend of AI that takes over every Nobel prize going forward? ๐Ÿค”

    Open Source LLMs & VLMs: The Community is COOKING!

    The open-source AI community consistently punches above its weight, and this week was no exception. We saw some truly impressive releases that deserve a standing ovation. First off, the TxT360 dataset (blog, dataset). Nisten, resident technical expert, broke down the immense effort: "The amount of DevOps andโ€ฆoperations to do this work is pretty rough."

    This globally deduplicated 15+ trillion-token corpus combines the best of Common Crawl with a curated selection of high-quality sources, setting a new standard for open-source LLM training. We talked about the importance of deduplication for model training - avoiding the "memorization" of repeated information that can skew a model's understanding of language. TxT360 takes a 360-degree approach to data quality and documentation โ€“ a huge win for accessibility.

    Apache 2 Multimodal MoE from Rhymes AI called Aria (blog, HF, Try It )

    Next, the Rhymes Aria model (25.3B total and only 3.9B active parameters!) This multimodal marvel operates as a Mixture of Experts (MoE), meaning it activates only the necessary parts of its vast network for a given task, making it surprisingly efficient. Aria excels in understanding image and video inputs, features a generous 64K token context window, and is available under the Apache 2 license โ€“ music to open-source developersโ€™ ears! We even discussed its coding capabilities: imagine pasting images of code and getting intelligent responses.

    I particularly love the focus on long multimodal input understanding (think longer videos) and super high resolution image support.

    I uploaded this simple pin-out diagram of RaspberriPy and it got all the right answers correct! Including ones I missed myself (and won against Gemini 002 and the new Reka Flash!)

    Big Companies and APIs

    OpenAI new Agentic benchmark, can it compete with MLEs on Kaggle?

    OpenAI snuck in a new benchmark, MLEBench (Paper, Github), specifically designed to evaluate AI agents performance on Machine Learning Engineering tasks. Designed around a curated collection of Kaggle competitions, creating a diverse set of challenging tasks that test real-world ML engineering skills such as training models, preparing datasets, and running experiments.

    They found that the best-performing setup--OpenAI's o1-preview with AIDE scaffolding--achieves at least the level of a Kaggle bronze medal in 16.9% of competitions (though there are some that throw shade on this score)

    Meta comes for our reality with Movie Gen

    But let's be honest, Meta stole the show this week with Movie Gen (blog). This isnโ€™t your average video generation model; itโ€™s like something straight out of science fiction. Imagine creating long, high-definition videos, with different aspect ratios, personalized elements, and accompanying audio โ€“ all from text and image prompts. It's like the Holodeck is finally within reach!

    Unfortunately, despite hinting at its size (30B) Meta is not releasing this model (just yet) nor is it available widely so far! But we'll keep our fingers crossed that it drops before SORA.

    One super notable thing is, this model generates audio as well to accompany the video and it's quite remarkable. We listened to a few examples from Metaโ€™s demo, and the sound effects were truly remarkable โ€“ everything from fireworks to rustling leaves. This model isn't just creating video, it's crafting experiences. (Sound on for the next example!)

    They also have personalization built in, which is showcased here by one of the leads of LLama ,Roshan, as a scientist doing experiments and the realism is quite awesome to see (but I get why they are afraid of releasing this in open weights)

    This Weekโ€™s Buzz: What I learned at Weights & Biases this week

    My "buzz" this week was less about groundbreaking models and more about mastering the AI tools we have. We had a team meeting to share our best tips and tricks for using Cursor, and when I shared those insights on X (thread), they went surprisingly viral!

    The big takeaway from the thread? Composer, Cursorโ€™s latest feature, is a true game-changer. It allows for more complex refactoring and code generation across multiple files โ€“ the kind of stuff that would take hours manually. If you haven't tried Composer, you're seriously missing out. We also covered strategies for leveraging different models for specific tasks, like using O1 mini for outlining and then switching to the more robust Cloud 3.5 for generating code. Another gem we uncovered: selecting any text in the console and hitting opt+D will immediately send it to the chat to debug, super useful!

    Over at Weights & Biases, my talented teammate, Soumik, released HEMM (X, Github), a comprehensive benchmark specifically designed for text-to-image generation models. Want to know how different models fare on image quality and prompt comprehension? Head over to the leaderboard on Weave (Leaderboard) and find out! And yes, it's true, Weave, our LLM observability tool, is multimodal (well within the theme of today's update)

    Voice and Audio: Real-Time Conversations and the Quest for Affordable AI

    OpenAI's DevDay was just a few weeks back, but the ripple effects of their announcements are still being felt. The big one for voice AI enthusiasts like myself? The RealTime API, offering developers a direct line to Advanced Voice Mode. My initial reaction was pure elation โ€“ finally, a chance to build some seriously interactive voice experiences that sound incredible and in near real time!

    That feeling was quickly followed by a sharp intake of breath when I saw the price tag. As I discovered building my Halloween project, real-time streaming of this caliber isnโ€™t exactly budget-friendly (yet!). Kwindla from trydaily.com, a voice AI expert, joined the show to shed some light on this issue.

    We talked about the challenges of scaling these models and the complexities of context management in real-time audio processing. The conversation shifted to how OpenAI's RealTime API isnโ€™t just about the model itself but also the innovative way they're managing the user experience and state within a conversation. He pointed out, however, that what we see and hear from the API isnโ€™t exactly whatโ€™s going on under the hood, โ€œWhat the model hears and what the transcription events give you back are not the sameโ€. Turns out, OpenAI relies on Whisper for generating text transcriptions โ€“ itโ€™s not directly from the voice model.

    The pricing really threw me though, only testing a little bit, not even doing anything on production, and OpenAI charged almost 10$, the same conversations are happening across Reddit and OpenAI forums as well.

    Hallo-Weave project update:

    So as I let folks know on the show, I'm building a halloween AI decoration as a project, and integrating it into Weights & Biases Weave (that's why it's called HalloWeave)

    After performing brain surgery, futzing with wires and LEDs, I finally have it set up so it wakes up on a trigger word (it's "Trick or Treat!"), takes a picture with the webcam (actual webcam, raspberryPi camera was god awful) and sends it to Gemini Flash to detect which costume this is and write a nice customized greeting.

    Then I send that text to Cartesia to generate the speech using a British voice, and then I play it via a bluetooth speaker. Here's a video of the last stage (which still had some bluetooth issues, it's a bit better now)

    Next up: I should decide if I care to integrate OpenAI Real time (and pay a LOT of $$$ for it) or fallback to existing LLM - TTS services and let kids actually have a conversation with the toy!

    Stay tuned for more updates as we get closer to halloween, the project is open source HERE and the Weave dashboard will be open once it's live.

    ThursdAI - Recaps of the most high signal AI weekly spaces is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

    One More Thingโ€ฆ UIThub!

    Before signing off, one super useful tool for you! It's so useful I recorded (and created an edit) video on it. I've also posted it on my brand new TikTok, Instagram, Youtube and Linkedin accounts, where it promptly did not receive any views, but hey, gotta start somewhere right? ๐Ÿ˜‚

    Phew! Thatโ€™s a wrap for this weekโ€™s ThursdAI. From Nobel Prizes to new open-source tools, and even meta's incredibly promising (but still locked down) video gen models, the world of AI continues to surprise and delight (and maybe cause a mild existential crisis or two!). I'd love to hear your thoughts โ€“ what caught your eye? Are you building anything cool? Let me know in the comments, and I'll see you back here next week for more AI adventures! Oh, and don't forget to subscribe to the podcast (five-star ratings always appreciated ๐Ÿ˜‰).



    This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
    1 hr 31 min
  • ๐Ÿ“† ThursdAI - Oct 3 - OpenAI RealTime API, ChatGPT Canvas & other DevDay news (how I met Sam Altman), Gemini 1.5 8B is basically free, BFL makes FLUX 1.1 6x faster, Rev breaks whisper records...

    Hey, it's Alex. Ok, so mind is officially blown. I was sure this week was going to be wild, but I didn't expect everyone else besides OpenAI to pile on, exactly on ThursdAI.

    Coming back from Dev Day (number 2) and am still processing, and wanted to actually do a recap by humans, not just the NotebookLM one I posted during the keynote itself (which was awesome and scary in a "will AI replace me as a podcaster" kind of way), and was incredible to have Simon Willison who was sitting just behind me most of Dev Day, join me for the recap!

    But then the news kept coming, OpenAI released Canvas, which is a whole new way of interacting with chatGPT, BFL released a new Flux version that's 8x faster, Rev released a Whisper killer ASR that does diarizaiton and Google released Gemini 1.5 Flash 8B, and said that with prompt caching (which OpenAI now also has, yay) this will cost a whopping 0.01 / Mtok. That's 1 cent per million tokens, for a multimodal model with 1 million context window. ๐Ÿคฏ

    This whole week was crazy, as last ThursdAI after finishing the newsletter I went to meet tons of folks at the AI Tinkerers in Seattle, and did a little EvalForge demo (which you can see here) and wanted to share EvalForge with you as well, it's early but very promising so feedback and PRs are welcome!

    WHAT A WEEK, TL;DR for those who want the links and let's dive in ๐Ÿ‘‡

    * OpenAI - Dev Day Recap (Alex, Simon Willison)

    * Recap of Dev Day

    * RealTime API launched

    * Prompt Caching launched

    * Model Distillation is the new finetune

    * Finetuning 4o with images (Skalski guide)

    * Fireside chat Q&A with Sam

    * Open Source LLMs

    * NVIDIA finally releases NVML (HF)

    * This weeks Buzz

    * Alex discussed his demo of EvalForge at the AI Tinkers event in Seattle in "This Week's Buzz". (Demo, EvalForge, AI TInkerers)

    * Big Companies & APIs

    * Google has released Gemini Flash 8B - 0.01 per million tokens cached (X, Blog)

    * Voice & Audio

    * Rev breaks SOTA on ASR with Rev ASR and Rev Diarize (Blog, Github, HF)

    * AI Art & Diffusion & 3D

    * BFL relases Flux1.1[pro] - 3x-6x faster than 1.0 and higher quality (was ๐Ÿซ) - (Blog, Try it)

    The day I met Sam Altman / Dev Day recap

    Last Dev Day (my coverage here) was a "singular" day in AI for me, given it also had the "keep AI open source" with Nous Research and Grimes, and this Dev Day I was delighted to find out that the vibe was completely different, and focused less on bombastic announcements or models, but on practical dev focused things.

    This meant that OpenAI cherry picked folks who actively develop with their tools, and they didn't invite traditional media, only folks like yours truly, @swyx from Latent space, Rowan from Rundown, Simon Willison and Dan Shipper, you know, newsletter and podcast folks who actually build!

    This also allowed for many many OpenAI employees who work on the products and APIs we get to use, were there to receive feedback, help folks with prompting, and just generally interact with the devs, and build that community. I want to shoutout my friends Ilan (who was in the keynote as the strawberry salesman interacting with RealTime API agent), Will DePue from the SORA team, with whom we had an incredible conversation about ethics and legality of projects, Christine McLeavey who runs the Audio team, with whom I shared a video of my daughter crying when chatGPT didn't understand her, Katia, Kevin and Romain on the incredible DevEx/DevRel team and finally, my new buddy Jason who does infra, and was fighting bugs all day and only joined the pub after shipping RealTime to all of us.

    I've collected all these folks in a convenient and super high signal X list here so definitely give that list a follow if you'd like to tap into their streams

    For the actual announcements, I've already covered this in my Dev Day post here (which was payed subscribers only, but is now open to all) and Simon did an incredible summary on his Substack as well

    The highlights were definitely the new RealTime API that let's developers build with Advanced Voice Mode, Prompt Caching that will happen automatically and reduce all your long context API calls by a whopping 50% and finetuning of models that they are rebranding into Distillation and adding new tools to make it easier (including Vision Finetuning for the first time!)

    Meeting Sam Altman

    While I didn't get a "media" pass or anything like this, and didn't really get to sit down with OpenAI execs (see Swyx on Latent Space for those conversations), I did have a chance to ask Sam multiple things.

    First at the closing fireside chat between Sam and Kevin Weil (CPO at OpenAI), Kevin first asked Sam a bunch of questions, and then they gave out the microphones to folks, and I asked the only question that got Sam to smile

    Sam and Kevin went on for a while, and that Q&A was actually very interesting, so much so, that I had to recruit my favorite Notebook LM podcast hosts, to go through it and give you an overview, so here's that Notebook LM, with the transcript of the whole Q&A (maybe i'll publish it as a standalone episode? LMK in the comments)

    After the official day was over, there was a reception, at the same gorgeous Fort Mason location, with drinks and light food, and as you might imagine, this was great for networking.

    But the real post dev day event was hosted by OpenAI devs at a bar, Palm House, which both Sam and Greg Brokman just came to and hung out with folks. I missed Sam last time and was very eager to go and ask him follow up questions this time, when I saw he was just chilling at that bar, talking to devs, as though he didn't "just" complete the largest funding round in VC history ($6.6B at $175B valuation) and went through a lot of drama/turmoil with the departure of a lot of senior leadership!

    Sam was awesome to briefly chat with, tho as you might imagine, it was loud and tons of folks wanted selfies, but we did discuss how AI affects the real world, job replacement stuff were brought up, and how developers are using the OpenAI products.

    What we learned, thanks to Sigil, is that o1 was named partly as a "reset" like the main blogpost claimed and partly as "alien of extraordinary ability" , which is the the official designation of the o1 visa, and that Sam came up with this joke himself.

    Is anyone here smarter than o1? Do you think you still will by o2?

    One of the highest impact questions was by Sam himself to the audience.

    Who feels like they've spent a lot of time with O1, and they would say, like, I feel definitively smarter than that thing?

    โ€” Sam Altman

    When Sam asked this at first, a few hands hesitatingly went up. He then followed up with

    Do you think you still will by O2? No one. No one taking the bet.One of the challenges that we face is like, we know how to go do this thing that we think will be like, at least probably smarter than all of us in like a broad array of tasks

    This was a very palpable moment that folks looked around and realized, what OpenAI folks have probably internalized a long time ago, we're living in INSANE times, and even those of us at the frontier or research, AI use and development, don't necessarily understand or internalize how WILD the upcoming few months, years will be.

    And then we all promptly forgot to have an existential crisis about it, and took our self driving Waymo's to meet Sam Altman at a bar ๐Ÿ˜‚

    This weeks Buzz from Weights & Biases

    Hey so... after finishing ThursdAI last week I went to Seattle Tinkerers event and gave a demo (and sponsored the event with a raffle of Meta Raybans). I demoed our project called EvalForge, which I built the frontend of and my collegue Anish on backend, as we tried to replicate the Who validates the validators paper by Shreya Shankar, hereโ€™s that demo, and EvalForge Github for many of you who asked to see it.

    Please let me know what you think, I love doing demos and would love feedback and ideas for the next one (coming up in October!)

    OpenAI chatGPT Canvas - a complete new way to interact with chatGPT

    Just 2 days after Dev Day, and as breaking news during the show, OpenAI also shipped a new way to interact with chatGPT, called Canvas!

    Get ready to say goodbye to simple chats and hello to a whole new era of AI collaboration! Canvas, a groundbreaking interface that transforms ChatGPT into a true creative partner for writing and coding projects. Imagine having a tireless copy editor, a brilliant code reviewer, and an endless source of inspiration all rolled into one โ€“ that's Canvas!

    Canvas moves beyond the limitations of a simple chat window, offering a dedicated space where you and ChatGPT can work side-by-side. Canvas opens in a separate window, allowing for a more visual and interactive workflow. You can directly edit text or code within Canvas, highlight sections for specific feedback, and even use a handy menu of shortcuts to request tasks like adjusting the length of your writing, debugging code, or adding final polish. And just like with your favorite design tools, you can easily restore previous versions using the back button.

    Per Karina, OpenAI has trained a special GPT-4o model specifically for Canvas, enabling it to understand the context of your project and provide more insightful assistance. They used synthetic data, generated by O1 which led them to outperform the basic version of GPT-4o by 30% in accuracy.

    A general pattern emerges, where new frontiers in intelligence are advancing also older models (and humans as well).

    Gemini Flash 8B makes intelligence essentially free

    Google folks were not about to take this week litely and decided to hit back with one of the most insane upgrades to pricing I've seen. The newly announced Gemini Flash 1.5 8B is goint to cost just... $0.01 per million tokens ๐Ÿคฏ (when using caching, 3 cents when not cached)

    This basically turns intelligence free. And while it is free, it's still their multimodal model (supports images) and has a HUGE context window of 1M tokens.

    The evals look ridiculous as well, this 8B param model, now almost matches Flash from May of this year, less than 6 month ago, while giving developers 2x the rate limits and lower latency as well.

    What will you do with free intelligence? What will you do with free intelligence of o1 quality in a year? what about o2 quality in 3 years?

    Bye Bye whisper? Rev open sources Reverb and Reverb Diarize + turbo models (Blog, HF, Github)

    With a "WTF just happened" breaking news, a company called Rev.com releases what they consider a SOTA ASR model, that obliterates Whisper (English only for now) on metrics like WER, and includes a specific diarization focused model.

    Trained on 200,000 hours of English speech, expertly transcribed by humans, which according to their claims, is the largest dataset that any ASR model has been trained on, they achieve some incredible results that blow whisper out of the water (lower WER is better)

    They also released a seemingly incredible diarization model, which helps understand who speaks when (and is usually added on top of Whisper)

    For diarization, Rev used the high-performance pyannote.audio library to fine-tune existing models on 26,000 hours of expertly labeled data, significantly improving their performance

    While this is for English only, getting a SOTA transcription model in the open, is remarkable. Rev opened up this model on HuggingFace with a non commercial license, so folks can play around (and distill?) it, while also making it available in their API for very cheap and also a self hosted solution in a docker container

    Black Forest Labs feeding up blueberries - new Flux 1.1[pro] is here (Blog, Try It)

    What is a ThursdAI without multiple SOTA advancements in all fields of AI? In an effort to prove this to be very true, the folks behind FLUX, revealed that the mysterious ๐Ÿซ model that was trending on some image comparison leaderboards is in fact a new version of Flux pro, specifically 1.1[pro]

    FLUX1.1 [pro] provides six times faster generation than its predecessor FLUX.1 [pro] while also improving image quality, prompt adherence, and diversity

    Just a bit over 2 month since the inital release, and proving that they are THE frontier lab for image diffusion models, folks at BLF are dropping a model that outperforms their previous one on users voting and quality, while being a much faster!

    They have partnered with Fal, Together, Replicate to disseminate this model (it's not on X quite yet) but are now also offering developers direct access to their own API and at a competitive pricing of just 4 cents per image generation (while being faster AND cheaper AND higher quality than the previous Flux ๐Ÿ˜ฎ) and you can try it out on Fal here

    Phew! What a whirlwind! Even I need a moment to catch my breath after that AI news tsunami. But donโ€™t worry, the conversation doesn't end here. I barely scratched the surface of these groundbreaking announcements, so dive into the podcast episode for the full scoop โ€“ Simon Willisonโ€™s insights on OpenAIโ€™s moves are pure gold, and Maxim LaBonne spills the tea on Liquid AI's audacious plan to dethrone transformers (yes, you read that right). And for those of you who prefer skimming, check out my Dev Day summary (open to all now). As always, hit me up in the comments with your thoughts. What are you most excited about? Are you building anything cool with these new tools? Let's keep the conversation going!Alex



    This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
    1 hr 46 min
  • OpenAI Dev Day 2024 keynote

    Hey, Alex here. Super quick, as Iโ€™m still attending Dev Day, but I didnโ€™t want to leave you hanging (if you're a paid subscriber!), I have decided to outsource my job and give the amazing podcasters of NoteBookLM the whole transcript of the opening keynote of OpenAI Dev Day.

    You can see a blog of everything they just posted here

    Hereโ€™s a summary of all what was announced:

    * Developer-Centric Approach: OpenAI consistently emphasized the importance of developers in their mission to build beneficial AGI. The speaker stated, "OpenAI's mission is to build AGI that benefits all of humanity, and developers are critical to that mission... we cannot do this without you."

    * Reasoning as a New Frontier: The introduction of the GPT-4 series, specifically the "O1" models, marks a significant step towards AI with advanced reasoning capabilities, going beyond the limitations of previous models like GPT-3.

    * Multimodal Capabilities: OpenAI is expanding the potential of AI applications by introducing multimodal capabilities, particularly focusing on real-time speech-to-speech interaction through the new Realtime API.

    * Customization and Fine-Tuning: Empowering developers to customize models is a key theme. OpenAI introduced Vision for fine-tuning with images and announced easier access to fine-tuning with model distillation tools.

    * Accessibility and Scalability: OpenAI demonstrated a commitment to making AI more accessible and cost-effective for developers through initiatives like price reductions, prompt caching, and model distillation tools.

    Important Ideas and Facts:

    1. The O1 Models:

    * Represent a shift towards AI models with enhanced reasoning capabilities, surpassing previous generations in problem-solving and logical thought processes.

    * O1 Preview is positioned as the most powerful reasoning model, designed for complex problems requiring extended thought processes.

    * O1 Mini offers a faster, cheaper, and smaller alternative, particularly suited for tasks like code debugging and agent-based applications.

    * Both models demonstrate advanced capabilities in coding, math, and scientific reasoning.

    * OpenAI highlighted the ability of O1 models to work with developers as "thought partners," understanding complex instructions and contributing to the development process.

    Quote: "The shift to reasoning introduces a new shape of AI capability. The ability for our model to scale and correct the process is pretty mind-blowing. So we are resetting the clock, and we are introducing a new series of models under the name O1."

    2. Realtime API:

    * Enables developers to build real-time AI experiences directly into their applications using WebSockets.

    * Launches with support for speech-to-speech interaction, leveraging the technology behind ChatGPT's advanced voice models.

    * Offers natural and seamless integration of voice capabilities, allowing for dynamic and interactive user experiences.

    * Showcased the potential to revolutionize human-computer interaction across various domains like driving, education, and accessibility.

    Quote: "You know, a lot of you have been asking about building amazing speech-to-speech experiences right into your apps. Well now, you can."

    3. Vision, Fine-Tuning, and Model Distillation:

    * Vision introduces the ability to use images for fine-tuning, enabling developers to enhance model performance in image understanding tasks.

    * Fine-tuning with Vision opens up opportunities in diverse fields such as product recommendations, medical imaging, and autonomous driving.

    * OpenAI emphasized the accessibility of these features, stating that "fine-tuning with Vision is available to every single developer."

    * Model distillation tools facilitate the creation of smaller, more efficient models by transferring knowledge from larger models like O1 and GPT-4.

    * This approach addresses cost concerns and makes advanced AI capabilities more accessible for a wider range of applications and developers.

    Quote: "With distillation, you take the outputs of a large model to supervise, to teach a smaller model. And so today, we are announcing our own model distillation tools."

    4. Cost Reduction and Accessibility:

    * OpenAI highlighted its commitment to lowering the cost of AI models, making them more accessible for diverse use cases.

    * Announced a 90% decrease in cost per token since the release of GPT-3, emphasizing continuous efforts to improve affordability.

    * Introduced prompt caching, automatically providing a 50% discount for input tokens the model has recently processed.

    * These initiatives aim to remove financial barriers and encourage wider adoption of AI technologies across various industries.

    Quote: "Every time we reduce the price, we see new types of applications, new types of use cases emerge. We're super far from the price equilibrium. In a way, models are still too expensive to be bought at massive scale."

    Conclusion:

    OpenAI DevDay conveyed a strong message of developer empowerment and a commitment to pushing the boundaries of AI capabilities. With new models like O1, the introduction of the Realtime API, and a dedicated focus on accessibility and customization, OpenAI is paving the way for a new wave of innovative and impactful AI applications developed by a global community.



    This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
    6 min
  • ๐Ÿ“… ThursdAI - Sep 26 - ๐Ÿ”ฅ Llama 3.2 multimodal & meta connect recap, new Gemini 002, Advanced Voice mode & more AI news

    Hey everyone, it's Alex (still traveling!), and oh boy, what a week again! Advanced Voice Mode is finally here from OpenAI, Google updated their Gemini models in a huge way and then Meta announced MultiModal LlaMas and on device mini Llamas (and we also got a "better"? multimodal from Allen AI called MOLMO!)

    From Weights & Biases perspective, our hackathon was a success this weekend, and then I went down to Menlo Park for my first Meta Connect conference, full of news and updates and will do a full recap here as well.

    ThursdAI - Recaps of the most high signal AI weekly spaces is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

    Overall another crazy week in AI, and it seems that everyone is trying to rush something out the door before OpenAI Dev Day next week (which I'll cover as well!) Get ready, folks, because Dev Day is going to be epic!

    TL;DR of all topics covered:

    * Open Source LLMs

    * Meta llama 3.2 Multimodal models (11B & 90B) (X, HF, try free)

    * Meta Llama 3.2 tiny models 1B & 3B parameters (X, Blog, download)

    * Allen AI releases MOLMO - open SOTA multimodal AI models (X, Blog, HF, Try It)

    * Big CO LLMs + APIs

    * OpenAI releases Advanced Voice Mode to all & Mira Murati leaves OpenAI

    * Google updates Gemini 1.5-Pro-002 and 1.5-Flash-002 (Blog)

    * This weeks Buzz

    * Our free course is LIVE - more than 3000 already started learning how to build advanced RAG++

    * Sponsoring tonights AI Tinkerers in Seattle, if you're in Seattle, come through for my demo

    * Voice & Audio

    * Meta also launches voice mode (demo)

    * Tools & Others

    * Project ORION - holographic glasses are here! (link)

    Meta gives us new LLaMas and AI hardware

    LLama 3.2 Multimodal 11B and 90B

    This was by far the biggest OpenSource release of this week (tho see below, may not be the "best"), as a rumored released finally came out, and Meta has given our Llama eyes! Coming with 2 versions (well 4 if you count the base models which they also released), these new MultiModal LLaMas were trained with an adapter architecture, keeping the underlying text models the same, and placing a vision encoder that was trained and finetuned separately on top.

    LLama 90B is among the best open-source mutlimodal models available

    โ€” Meta team at launch

    These new vision adapters were trained on a massive 6 Billion images, including synthetic data generation by 405B for questions/captions, and finetuned with a subset of 600M high quality image pairs.

    Unlike the rest of their models, the Meta team did NOT claim SOTA on these models, and the benchmarks are very good but not the best we've seen (Qwen 2 VL from a couple of weeks ago, and MOLMO from today beat it on several benchmarks)

    With text-only inputs, the Llama 3.2 Vision models are functionally the same as the Llama 3.1 Text models; this allows the Llama 3.2 Vision models to be a drop-in replacement for Llama 3.1 8B/70B with added image understanding capabilities.

    Seems like these models don't support multi image or video as well (unlike Pixtral for example) nor tool use with images.

    Meta will also release these models on meta.ai and every other platform, and they cited a crazy 500 million monthly active users of their AI services across all their apps ๐Ÿคฏ which marks them as the leading AI services provider in the world now.

    Llama 3.2 Lightweight Models (1B/3B)

    The additional and maybe more exciting thing that we got form Meta was the introduction of the small/lightweight models of 1B and 3B parameters.

    Trained on up to 9T tokens, and distilled / pruned from larger models, these are aimed for on-device inference (and by device here we mean from laptops to mobiles to soon... glasses? more on this later)

    In fact, meta released an IOS demo, that runs these models, takes a group chat, summarizes and calls the calendar tool to schedule based on the conversation, and all this happens on device without the info leaving to a larger model.

    They have also been able to prune down the LLama-guard safety model they released to under 500Mb and have had demos of it running on client side and hiding user input on the fly as the user types something bad!

    Interestingly, here too, the models were not SOTA, even in small category, with tiny models like Qwen 2.5 3B beating these models on many benchmarks, but they are outlining a new distillation / pruning era for Meta as they aim for these models to run on device, eventually even glasses (and some said Smart Thermostats)

    In fact they are so tiny, that the communtiy quantized them, released and I was able to download these models, all while the keynote was still going! Here I am running the Llama 3B during the developer keynote!

    Speaking AI - not only from OpenAI

    Zuck also showcased a voice based Llama that's coming to Meta AI (unlike OpenAI it's likely a pipeline of TTS/STT) but it worked really fast and Zuck was able to interrupt it.

    And they also showed a crazy animated AI avatar of a creator, that was fully backed by Llama, while the human creator was on stage, Zuck chatted with his avatar and reaction times were really really impressive.

    AI Hardware was glasses all along?

    Look we've all seen the blunders of this year, the Humane AI Ping, the Rabbit R1 (which sits on my desk and I haven't recharged in two months) but maybe Meta is the answer here?

    Zuck took a bold claim that glasses are actually the perfect form factor for AI, it sits on your face, sees what you see and hears what you hear, and can whisper in your ear without disrupting the connection between you and your conversation partner.

    They haven't announced new Meta Raybans, but did update the lineup with a new set of transition lenses (to be able to wear those glasses inside and out) and a special edition clear case pair that looks very sleek + new AI features like memories to be able to ask the glasses "hey Meta where did I park" or be able to continue the conversation. I had to get me a pair of this limited edition ones!

    Project ORION - first holographic glasses

    And of course, the biggest announcement of the Meta Connect was the super secret decade old project of fully holographic AR glasses, which they called ORION.

    Zuck introduced these as the most innovative and technologically dense set of glasses in the world. They always said the form factor will become just "glasses" and they actually did it ( a week after Snap spectacles ) tho those are not going to get released to any one any time soon, hell they only made a few thousand of these and they are extremely expensive.

    With 70 deg FOV, cameras, speakers and a compute puck, these glasses pack a full day battery with under 100grams of weight, and have a custom silicon, custom displays with MicroLED projector and just... tons of more innovation in there.

    They also come in 3 pieces, the glasses themselves, the compute wireless pack that will hold the LLaMas in your pocket and the EMG wristband that allows you to control these devices using muscle signals.

    These won't ship as a product tho so don't expect to get them soon, but they are real, and will allow Meta to build the product that we will get on top of these by 2030

    AI usecases

    So what will these glasses be able to do? well, they showed off a live translation feature on stage that mostly worked, where you just talk and listen to another language in near real time, which was great. There are a bunch of mixed reality games, you'd be able to call people and see them in your glasses on a virtual screen and soon you'll show up as an avatar there as well.

    The AI use-case they showed beyond just translation was MultiModality stuff, where they had a bunch of ingredients for a shake, and you could ask your AI assistant, which shake you can make with what it sees. Do you really need

    I'm so excited about these to finally come to people I screamed in the audience ๐Ÿ‘€๐Ÿ‘“

    OpenAI gives everyone* advanced voice mode

    It's finally here, and if you're paying for chatGPT you know this, the long announced Advanced Voice Mode for chatGPT is now rolled out to all plus members.

    The new updated since the beta are, 5 new voices (Maple, Spruce, Vale, Arbor and Sol), finally access to custom instructions and memory, so you can ask it to remember things and also to know who you are and your preferences (try saving your jailbreaks there)

    Unfortunately, as predicted, by the time it rolled out to everyone, this feels way less exciting than it did 6 month ago, the model is way less emotional, refuses to sing (tho folks are making it anyway) and generally feels way less "wow" than what we saw. Less "HER" than we wanted for sure Seriously, they nerfed the singing! Why OpenAI, why?

    Pro tip of mine that went viral : you can set your action button on the newer iphones to immediately start the voice conversation with 1 click.

    *This new mode is not available in EU

    This weeks Buzz - our new advanced RAG++ course is live

    I had an awesome time with my colleagues Ayush and Bharat today, after they finally released a FREE advanced RAG course they've been working so hard on for the past few months! Definitely check out our conversation, but better yet, why don't you roll into the course? it's FREE and you'll get to learn about data ingestion, evaluation, query enhancement and more!

    New Gemini 002 is 50% cheaper, 2x faster and better at MMLU-pro

    It seems that every major lab (besides Anthropic) released a big thing this week to try and get under Meta's skin?

    Google announced an update to their Gemini Pro/Flash models, called 002, which is a very significant update!

    Not only are these models 50% cheaper now (Pro price went down by 50% on <128K context lengths), they are 2x faster on outputs with 3x lower latency on first tokens. It's really quite something to see

    The new models have also improved scores, with the Flash models (the super cheap ones, remember) from September, now coming close to or beating the Pro scores from May 2024!

    Definitely a worthy update from the team at Google!

    Hot off the press, the folks at Google Labs also added a feature to the awesome NotebookLM that allows it to summarize over 50h of youtube videos in the crazy high quality Audio Overview feature!

    ThursdAI - Recaps of the most high signal AI weekly spaces is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

    That's it for the week, we of course chatted about way way more during the show, so make sure to listen to the podcast this week, but otherwise, signing off for this week, as I travel back home for a weekend, before returning to SF for the OpenAI dev day next week!

    Expect full Dev Day coverage live next tuesday and a recap on the newsletter.

    Meanwhile, if you've already subscribed, please share this newsletter with 1 or two people who are interested in AI ๐Ÿ™‡โ€โ™‚๏ธ and see you next week.



    This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
    1 hr 48 min
  • ThursdAI - Sep 19 - ๐Ÿ‘‘ Qwen 2.5 new OSS king LLM, MSFT new MoE, Nous Research's Forge announcement, and Talking AIs in the open source!

    Hey folks, Alex here, back with another ThursdAI recap โ€“ and let me tell you, this week's episode was a whirlwind of open-source goodness, mind-bending inference techniques, and a whole lotta talk about talking AIs! We dove deep into the world of LLMs, from Alibaba's massive Qwen 2.5 drop to the quirky, real-time reactions of Moshi.

    We even got a sneak peek at Nous Research's ambitious new project, Forge, which promises to unlock some serious LLM potential. So grab your pumpkin spice latte (it's that time again isn't it? ๐Ÿ) settle in, and let's recap the AI awesomeness that went down on ThursdAI, September 19th!

    ThursdAI is brought to you (as always) by Weights & Biases, we still have a few spots left in our Hackathon this weekend and our new advanced RAG course is now released and is FREE to sign up!

    TL;DR of all topics + show notes and links

    * Open Source LLMs

    * Alibaba Qwen 2.5 models drop + Qwen 2.5 Math and Qwen 2.5 Code (X, HF, Blog, Try It)

    * Qwen 2.5 Coder 1.5B is running on a 4 year old phone (Nisten)

    * KyutAI open sources Moshi & Mimi (Moshiko & Moshika) - end to end voice chat model (X, HF, Paper)

    * Microsoft releases GRIN-MoE - tiny (6.6B active) MoE with 79.4 MMLU (X, HF, GIthub)

    * Nvidia - announces NVLM 1.0 - frontier class multimodal LLMS (no weights yet, X)

    * Big CO LLMs + APIs

    * OpenAI O1 results from LMsys do NOT disappoint - vibe checks also confirm, new KING llm in town (Thread)

    * NousResearch announces Forge in waitlist - their MCTS enabled inference product (X)

    * This weeks Buzz - everything Weights & Biases related this week

    * Judgement Day (hackathon) is in 2 days! Still places to come hack with us Sign up

    * Our new RAG Course is live - learn all about advanced RAG from WandB, Cohere and Weaviate (sign up for free)

    * Vision & Video

    * Youtube announces DreamScreen - generative AI image and video in youtube shorts ( Blog)

    * CogVideoX-5B-I2V - leading open source img2video model (X, HF)

    * Runway, DreamMachine & Kling all announce text-2-video over API (Runway, DreamMachine)

    * Runway announces video 2 video model (X)

    * Tools

    * Snap announces their XR glasses - have hand tracking and AI features (X)

    Open Source Explosion!

    ๐Ÿ‘‘ Qwen 2.5: new king of OSS llm models with 12 model releases, including instruct, math and coder versions

    This week's open-source highlight was undoubtedly the release of Alibaba's Qwen 2.5 models. We had Justin Lin from the Qwen team join us live to break down this monster drop, which includes a whopping seven different sizes, ranging from a nimble 0.5B parameter model all the way up to a colossal 72B beast! And as if that wasn't enough, they also dropped Qwen 2.5 Coder and Qwen 2.5 Math models, further specializing their LLM arsenal. As Justin mentioned, they heard the community's calls for 14B and 32B models loud and clear โ€“ and they delivered! "We do not have enough GPUs to train the models," Justin admitted, "but there are a lot of voices in the community...so we endeavor for it and bring them to you." Talk about listening to your users!

    Trained on an astronomical 18 trillion tokens (thatโ€™s even more than Llama 3.1 at 15T!), Qwen 2.5 shows significant improvements across the board, especially in coding and math. They even open-sourced the previously closed-weight Qwen 2 VL 72B, giving us access to the best open-source vision language models out there. With a 128K context window, these models are ready to tackle some serious tasks. As Nisten exclaimed after putting the 32B model through its paces, "It's really practicalโ€ฆI was dumping in my docs and my code base and then like actually asking questions."

    It's safe to say that Qwen 2.5 coder is now the best coding LLM that you can use, and just in time for our chat, a new update from ZeroEval confirms, Qwen 2.5 models are the absolute kings of OSS LLMS, beating Mistral large, 4o-mini, Gemini Flash and other huge models with just 72B parameters ๐Ÿ‘

    Moshi: The Chatty Cathy of AI

    We've covered Moshi Voice back in July, and they have promised to open source the whole stack, and now finally they did! Including the LLM and the Mimi Audio Encoder!

    This quirky little 7.6B parameter model is a speech-to-speech marvel, capable of understanding your voice and responding in kind. It's an end-to-end model, meaning it handles the entire speech-to-speech process internally, without relying on separate speech-to-text and text-to-speech models.

    While it might not be a logic genius, Moshi's real-time reactions are undeniably uncanny. Wolfram Ravenwolf described the experience: "It's uncanny when you don't even realize you finished speaking and it already starts to answer." The speed comes from the integrated architecture and efficient codecs, boasting a theoretical response time of just 160 milliseconds!

    Moshi uses (also open sourced) Mimi neural audio codec, and achieves 12.5 Hz representation with just 1.1 kbps bandwidth.

    You can download it and run on your own machine or give it a try here just don't expect a masterful conversationalist hehe

    Gradient-Informed MoE (GRIN-MoE): A Tiny Titan

    Just before our live show, Microsoft dropped a paper on GrinMoE, a gradient-informed Mixture of Experts model. We were lucky enough to have the lead author, Liyuan Liu (aka Lucas), join us impromptu to discuss this exciting development. Despite having only 6.6B active parameters (16 x 3.8B experts), GrinMoE manages to achieve remarkable performance, even outperforming larger models like Phi-3 on certain benchmarks. It's a testament to the power of clever architecture and training techniques. Plus, it's open-sourced under the MIT license, making it a valuable resource for the community.

    NVIDIA NVLM: A Teaser for Now

    NVIDIA announced NVLM 1.0, their own set of multimodal LLMs, but alas, no weights were released. Weโ€™ll have to wait and see how they stack up against the competition once they finally let us get our hands on them. Interestingly, while claiming SOTA on some vision tasks, they haven't actually compared themselves to Qwen 2 VL, which we know is really really good at vision tasks ๐Ÿค”

    Nous Research Unveils Forge: Inference Time Compute Powerhouse (beating o1 at AIME Eval!)

    Fresh off their NousCon event, Karan and Shannon from Nous Research joined us to discuss their latest project, Forge. Described by Shannon as "Jarvis on the front end," Forge is an inference engine designed to push the limits of whatโ€™s possible with existing LLMs. Their secret weapon? Inference-time compute. By implementing sophisticated techniques like Monte Carlo Tree Search (MCTS), Forge can outperform larger models on complex reasoning tasks beating OpenAI's o1-preview at the AIME Eval, competition math benchmark, even with smaller, locally runnable models like Hermes 70B. As Karan emphasized, โ€œWeโ€™re actually just scoring with Hermes 3.1, which is available to everyone already...we can scale it up to outperform everything on math, just using a system like this.โ€

    Forge isn't just about raw performance, though. It's built with usability and transparency in mind. Unlike OpenAI's 01, which obfuscates its chain of thought reasoning, Forge provides users with a clear visual representation of the model's thought process. "You will still have access in the sidebar to the full chain of thought," Shannon explained, adding, โ€œThereโ€™s a little visualizer and it will show you the trajectory through the treeโ€ฆ youโ€™ll be able to see exactly what the model was doing and why the node was selected.โ€ Forge also boasts built-in memory, a graph database, and even code interpreter capabilities, initially supporting Python, making it a powerful platform for building complex LLM applications.

    Forge is currently in a closed beta, but a waitlist is open for eager users. Karan and Shannon are taking a cautious approach to the rollout, as this is Nous Researchโ€™s first foray into hosting a product. For those lucky enough to gain access, Forge offers a tantalizing glimpse into the future of LLM interaction, promising greater transparency, improved reasoning, and more control over the model's behavior.

    For ThursdAI readers early, here's a waitlist form to test it out!

    Big Companies and APIs: The Reasoning Revolution

    OpenAIโ€™s 01: A New Era of LLM Reasoning

    The big story in the Big Tech world is OpenAI's 01. Since we covered it live last week as it dropped, many of us have been playing with these new reasoning models, and collecting "vibes" from the community. These models represent a major leap in reasoning capabilities, and the results speak for themselves.

    01 Preview claimed the top spot across the board on the LMSys Arena leaderboard, demonstrating significant improvements in complex tasks like competition math and coding. Even the smaller 01 Mini showed impressive performance, outshining larger models in certain technical areas. (and the jump in ELO score above the rest in MATH is just incredible to see!) and some folks made this video viral, of a PHD candidate reacting to 01 writing in 1 shot, code that took him a year to write, check it out, itโ€™s priceless.

    One key aspect of 01 is the concept of โ€œinference-time computeโ€. As Noam Brown from OpenAI calls it, this represents a "new scaling paradigm", allowing the model to spend more time โ€œthinkingโ€ during inference, leading to significantly improved performance on reasoning tasks. The implications of this are vast, opening up the possibility of LLMs tackling long-horizon problems in areas like drug discovery and physics.

    However, the opacity surrounding 01โ€™s chain of thought reasoning being hidden/obfuscated and the ban on users asking about it was a major point of contention at least within the ThursdAI chat. As Wolfram Ravenwolf put it, "The AI gives you an answer and you can't even ask how it got there. That is the wrong direction." as he was referring to the fact that not only is asking about the reasoning impossible, some folks were actually getting threatening emails and getting banned from using the product all together ๐Ÿ˜ฎ

    This Week's Buzz: Hackathons and RAG Courses!

    We're almost ready to host our Weights & Biases Judgment Day Hackathon (LLMs as a judge, anyone?) with a few spots left, so if you're reading this and in SF, come hang out with us!

    And the main thing I gave an update about is our Advanced RAG course, packed with insights from experts at Weights & Biases, Cohere, and Weaviate. Definitely check those out if you want to level up your LLM skills (and it's FREE in our courses academy!)

    Vision & Video: The Rise of Generative Video

    Generative video is having its moment, with a flurry of exciting announcements this week. First up, the open-source CogVideoX-5B-I2V, which brings accessible image-to-video capabilities to the masses. It's not perfect, but being able to generate video on your own hardware is a game-changer.

    On the closed-source front, YouTube announced the integration of generative AI into YouTube Shorts with their DreamScreen feature, bringing AI-powered video generation to a massive audience. We also saw API releases from three leading video model providers: Runway, DreamMachine, and Kling, making it easier than ever to integrate generative video into applications. Runway even unveiled a video-to-video model, offering even more control over the creative process, and it's wild, check out what folks are doing with video-2-video!

    One last thing here, Kling is adding a motion brush feature to help users guide their video generations, and it just looks so awesome I wanted to show you

    Whew! That was one hell of a week, tho from the big companies perspective, it was a very slow week, getting a new OSS king, an end to end voice model and a new hint of inference platform from Nous, and having all those folks come to the show was awesome!

    If you're reading all the way down to here, it seems that you like this content, why not share it with 1 or two friends? ๐Ÿ‘‡ And as always, thank you for reading and subscribing! ๐Ÿซถ

    P.S - Iโ€™m traveling for the next two weeks, and this week the live show was live recorded from San Francisco, thanks to my dear friends swyx & Alessio for hosting my again in their awesome Latent Space pod studio at Solaris SF!



    This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
    1 hr 57 min
  • ๐Ÿ”ฅ ๐Ÿ“… ThursdAI - Sep 12 - OpenAI's ๐Ÿ“ is called 01 and is HERE, reflecting on Reflection 70B, Google's new auto podcasts & more AI news from last week

    March 14th, 2023 was the day ThursdAI was born, it was also the day OpenAI released GPT-4, and I jumped into a Twitter space and started chaotically reacting together with other folks about what a new release of a paradigm shifting model from OpenAI means, what are the details, the new capabilities. Today, it happened again!

    Hey, it's Alex, I'm back from my mini vacation (pic after the signature) and boy am I glad I decided to not miss September 12th! The long rumored ๐Ÿ“ thinking model from OpenAI, dropped as breaking news in the middle of ThursdAI live show, giving us plenty of time to react live!

    But before this, we already had an amazing show with some great guests! Devendra Chaplot from Mistral came on and talked about their newly torrented (yeah they did that again) Pixtral VLM, their first multi modal! , and then I had the honor to host Steven Johnson and Raiza Martin from NotebookLM team at Google Labs which shipped something so uncannily good, that I legit said "holy fu*k" on X in a reaction!

    So let's get into it (TL;DR and links will be at the end of this newsletter)

    OpenAI o1, o1 preview and o1-mini, a series of new "reasoning" models

    This is it folks, the strawberries have bloomed, and we finally get to taste them. OpenAI has released (without a waitlist, 100% rollout!) o1-preview and o1-mini models to chatGPT and API (tho only for tier-5 customers) ๐Ÿ‘ and are working on releasing 01 as well.

    These are models that think before they speak, and have been trained to imitate "system 2" thinking, and integrate chain-of-thought reasoning internally, using Reinforcement Learning and special thinking tokens, which allows them to actually review what they are about to say before they are saying it, achieving remarkable results on logic based questions.

    Specifically you can see the jumps in the very very hard things like competition math and competition code, because those usually require a lot of reasoning, which is what these models were trained to do well.

    New scaling paradigm

    Noam Brown from OpenAI calls this a "new scaling paradigm" and Dr Jim Fan explains why, with this new way of "reasoning", the longer the model thinks - the better it does on reasoning tasks, they call this "test-time compute" or "inference-time compute" as opposed to compute that was used to train the model. This shifting of computation down to inference time is the essence of the paradigm shift, as in, pre-training can be very limiting computationally as the models scale in size of parameters, they can only go so big until you have to start building out a huge new supercluster of GPUs to host the next training run (Remember Elon's Colossus from last week?).

    The interesting thing to consider here is, while current "thinking" times are ranging between a few seconds to a minute, imagine giving this model hours, days, weeks to think about new drug problems, physics problems ๐Ÿคฏ.

    Prompting o1

    Interestingly, a new prompting paradigm has also been introduced. These models now have CoT (think "step by step") built-in, so you no longer have to include it in your prompts. By simply switching to o1-mini, most users will see better results right off the bat. OpenAI has worked with the Devin team to test drive these models, and these folks found that asking the new models to just give the final answer often works better and avoids redundancy in instructions.

    The community of course will learn what works and doesn't in the next few hours, days, weeks, which is why we got 01-preview and not the actual (much better) o1.

    Safety implications and future plans

    According to Greg Brokman, this inference time compute also greatly helps with aligning the model to policies, giving it time to think about policies at length, and improving security and jailbreak preventions, not only logic.

    The folks at OpenAI are so proud of all of the above that they have decided to restart the count and call this series o1, but they did mention that they are going to release GPT series models as well, adding to the confusing marketing around their models.

    Open Source LLMs

    Reflecting on Reflection 70B

    Last week, Reflection 70B was supposed to launch live on the ThursdAI show, and while it didn't happen live, I did add it in post editing, and sent the newsletter, and packed my bag, and flew for my vacation. I got many DMs since then, and at some point couldn't resist checking and what I saw was complete chaos, and despite this, I tried to disconnect still until last night.

    So here's what I could gather since last night. The claims of a llama 3.1 70B finetune that Matt Shumer and Sahil Chaudhary from Glaive beating Sonnet 3.5 are proven false, nobody was able to reproduce those evals they posted and boasted about, which is a damn shame.

    Not only that, multiple trusted folks from our community, like Kyle Corbitt, Alex Atallah have reached out to Matt in to try to and get to the bottom of how such a thing would happen, and how claims like these could have been made in good faith. (or was there foul play)

    The core idea of something like Reflection is actually very interesting, but alas, the inability to replicate, but also to stop engaging with he community openly (I've reached out to Matt and given him the opportunity to come to the show and address the topic, he did not reply), keep the model on hugging face where it's still trending, claiming to be the world's number 1 open source model, all these smell really bad, despite multiple efforts on out part to give the benefit of the doubt here.

    As for my part in building the hype on this (last week's issues till claims that this model is top open source model), I addressed it in the beginning of the show, but then twitter spaces crashed, but unfortunately as much as I'd like to be able to personally check every thing I cover, I often have to rely on the reputation of my sources, which is easier with established big companies, and this time this approached failed me.

    This weeks Buzzzzzz - One last week till our hackathon!

    Look at this point, if you read this newsletter and don't know about our hackathon, then I really didn't do my job prompting it, but it's coming up, September 21-22 ! Join us, it's going to be a LOT of fun!

    ๐Ÿ–ผ๏ธ Pixtral 12B from Mistral

    Mistral AI burst onto the scene with Pixtral, their first multimodal model! Devendra Chaplot, research scientist at Mistral, joined ThursdAI to explain their unique approach, ditching fixed image resolutions and training a vision encoder from scratch.

    "We designed this from the ground up to...get the most value per flop," Devendra explained. Pixtral handles multiple images interleaved with text within a 128k context window - a far cry from the single-image capabilities of most open-source multimodal models. And to make the community erupt in thunderous applause (cue the clap emojis!) they released the 12 billion parameter model under the ultra-permissive Apache 2.0 license. You can give Pixtral a whirl on Hyperbolic, HuggingFace, or directly through Mistral.

    DeepSeek 2.5: When Intelligence Per Watt is King

    Deepseek 2.5 launched amid the reflection news and did NOT get the deserved attention it.... deserves. It folded (no deprecated) Deepseek Coder into 2.5 and shows incredible metrics and a truly next-gen architecture. "It's like a higher order MOE", Nisten revealed, "which has this whole like pile of brain and it just like picks every time, from that." ๐Ÿคฏ. DeepSeek 2.5 achieves maximum "intelligence per active parameter"

    Google's turning text into AI podcast for auditory learners with Audio Overviews

    Today I had the awesome pleasure of chatting with Steven Johnson and Raiza Martin from the NotebookLM team at Google Labs. NotebookLM is a research tool, that if you haven't used, you should definitely give it a spin, and this week they launched something I saw in preview and was looking forward to checking out and honestly was jaw-droppingly impressed today.

    NotebookLM allows you to upload up to 50 "sources" which can be PDFs, web links that they will scrape for you, documents etc' (no multimodality so far) and will allow you to chat with them, create study guides, dive deeper and add notes as you study.

    This week's update allows someone who doesn't like reading, to turn all those sources into a legit 5-10 minute podcast, and that sounds so realistic, that I was honestly blown away. I uploaded a documentation of fastHTML in there.. and well hear for yourself

    The conversation with Steven and Raiza was really fun, podcast definitely give it a listen!

    Not to mention that Google released (under waitlist) another podcast creating tool called illuminate, that will convert ArXiv papers into similar sounding very realistic 6-10 minute podcasts!

    ThursdAI - Recaps of the most high signal AI weekly spaces is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

    There are many more updates from this week, there was a whole Apple keynote I missed, which had a new point and describe feature with AI on the new iPhones and Apple Intelligence, Google also released new DataGemma 27B, and more things in TL'DR which are posted here in raw format

    See you next week ๐Ÿซก Thank you for being a subscriber, weeks like this are the reason we keep doing this! ๐Ÿ”ฅ Hope you enjoy these models, leave in comments what you think about them

    TL;DR in raw format

    * Open Source LLMs

    * Reflect on Reflection 70B & Matt Shumer (X, Sahil)

    * Mixtral releases Pixtral 12B - multimodal model (X, try it)

    * Pixtral is really good at OCR says swyx

    * Interview with Devendra Chaplot on ThursdAI

    * Initial reports of Pixtral beating GPT-4 on WildVision arena from AllenAI

    * JinaIA reader-lm-0.5b and reader-lm-1.5b (X)

    * ZeroEval updates

    * Deepseek 2.5 -

    * Deepseek coder is now folded into DeepSeek v2.5

    * 89 HumanEval (up from 84 from deepseek v2)

    * 9 on MT-bench

    * Google - DataGemma 27B (RIG/RAG) for improving results

    * Retrieval-Interleaved Generation

    * ๐Ÿค– DataGemma: AI models that connect LLMs to Google's Data Commons

    * ๐Ÿ“Š Data Commons: A vast repository of trustworthy public data

    * ๐Ÿ” Tackling AI hallucination by grounding LLMs in real-world data

    * ๐Ÿ” Two approaches: RIG (Retrieval-Interleaved Generation) and RAG (Retrieval-Augmented Generation)

    * ๐Ÿ” Preliminary results show enhanced accuracy and reduced hallucinations

    * ๐Ÿ”“ Making DataGemma open models to enable broader adoption

    * ๐ŸŒ Empowering informed decisions and deeper understanding of the world

    * ๐Ÿ” Ongoing research to refine the methodologies and scale the work

    * ๐Ÿ” Integrating DataGemma into Gemma and Gemini AI models

    * ๐Ÿค Collaborating with researchers and developers through quickstart notebooks

    * Big CO LLMs + APIs

    * Apple event

    * Apple Intelligence - launching soon

    * Visual Intelligence with a dedicated button

    * Google Illuminate - generate arXiv paper into multiple speaker podcasts (Website)

    * 5-10 min podcasts

    * multiple speakers

    * any paper

    * waitlist

    * has samples

    * sounds super cool

    * Google NotebookLM is finally available - multi modal research tool + podcast (NotebookLM)

    * Has RAG like abilities, can add sources from drive or direct web links

    * Currently not multimodal

    * Generation of multi speaker conversation about this topic to present it, sounds really really realistic

    * Chat with Steven and Raiza

    * OpenAI reveals new o1 models, and launches o1 preview and o1-mini in chat and API (X, Blog)

    * Trained with RL to think before it speaks with special thinking tokens (that you pay for)

    * new scaling paradigm

    * This weeks Buzz

    * Vision & Video

    * Adobe announces Firefly video model (X)

    * Voice & Audio

    * Hume launches EVI 2 (X)

    * Fish Speech 1.4 (X)

    * Instant Voice Cloning

    * Ultra low latenc

    * ~1GB model weights

    * LLaMA-Omni, a new model for speech interaction (X)

    * Tools

    * New Jina reader (X)



    This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
    1 hr 59 min

About ThursdAI - The top AI news from the past week

From the publisher's feed

Every ThursdAI, Alex Volkov hosts a panel of experts, ai engineers, data scientists and prompt spellcasters on twitter spaces, as we discuss everything major and important that happened in the world of AI for the past week.

More shows like ThursdAI - The top AI news from the past week

This Week in Startups by Jason Calacanis

This Week in Startups

1,292 Listeners

The Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch by Harry Stebbings

The Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch

541 Listeners

The a16z Show by Andreessen Horowitz

The a16z Show

1,089 Listeners

Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

Super Data Science: ML & AI Podcast with Jon Krohn

304 Listeners

Y Combinator Startup Podcast by Y Combinator

Y Combinator Startup Podcast

225 Listeners

Practical AI by Daniel Whitenack and Chris Benson

Practical AI

201 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

573 Listeners

Big Technology Podcast by Alex Kantrowitz

Big Technology Podcast

508 Listeners

The Artificial Intelligence Show by Paul Roetzer and Mike Kaput

The Artificial Intelligence Show

208 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

140 Listeners

Latent Space: The AI Engineer Podcast by Latent.Space

Latent Space: The AI Engineer Podcast

101 Listeners

This Day in AI Podcast by Michael Sharkey, Chris Sharkey

This Day in AI Podcast

222 Listeners

The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

The AI Daily Brief: Artificial Intelligence News and Analysis

682 Listeners

Everyday AI Podcast โ€“ An AI and ChatGPT Podcast by Everyday AI

Everyday AI Podcast โ€“ An AI and ChatGPT Podcast

109 Listeners

How I AI by Claire Vo

How I AI

157 Listeners