
Sign up to save your podcasts
Or


Wow what a week. I think I’ve reached to a level that I’m not phased by incredible weeks or days that happen in AI, but I… guess I still have much to learn!
TL;DR of everything we covered (aka Show Notes)
* Open Source LLMs
* Mixtral MoE - 8X7B experts dropped with a magnet link again (Announcement, HF, Try it)
* Mistral 0.2 instruct (Announcement, HF)
* Upstage Solar 10B - Tops the HF leaderboards (Announcement)
* Together -Striped Hyena architecture and new models (Announcement)
* EAGLE - a new decoding method for LLMs (Announcement, Github)
* Deci.ai - new SOTA 7B model
* Phi 2.0 weights are available finally from Microsoft (HF)
* QuiP - LLM quantization & Compression (link)
* Big CO LLMs + APIs
* Gemini Pro access over API (Announcement, Thread)
* Uses character pricing not token
* Mistral releases API inference server - La Platforme (API docs)
* Together undercuts Mistral with serving Mixtral by 70% and announces OAI compatible API
* OpenAI is open sourcing again - Releasing Weak-2-strong generalization paper and github! (announcement)
* Vision
* Gemini Pro api has vision AND video capabilities (API docs)
* AI Art & Diffusion
* Stability announces Zero123 - Zero Shot image to 3d model (Thread)
* Imagen 2 from google (link)
* Tools & Other
* Optimus from Tesla is coming, and it looks incredible
This week started on Friday, as we saw one of the crazier single days in the history of OSS AI that I can remember, and I’ve been doing this now for .. jesus, 9 months!
In a single say, we saw a new Mistral model release called Mixtral, which is a Mixture of Experts (like GPT4 is rumored to be) of 8x7B Mistrals, and beats GPT3.5, we saw a completely new architecture that competes with Transformers called HYENA from Tri Dao and Together.xyz + 2 new models trained with that architecture, we saw a new SoTA 2-bit quantization method called QuiP from cornell AND a new 3x faster decoding method for showing tokens to users after an LLM has done “thinking”.
And the best thing? All those advancements are stackable! What a day!
Then I went to NeurIPS2023 (which is where I am right now, writing these words!), which I cover at length at the second part of the podcast, but figured I’d write about it here as well, since it was such a crazy experience.
NeurIPS is the biggest AIML conference, I think they estimated 15K people from all over the world attending! Of course this brings many companies to sponsor, raise booths, give out swag and try to record!
Of course with my new position at Weights & Biases I had to come as well and experience this for myself!
Many of the attendees are customers of ours, and I was not expecting this amount of love given, just an incredible stream of people coming up to the booth, and saying how much they love the product!
So I manned the booth, did interviews and live streams, and connected with a LOT of folks and I gotta say, this whole NeurIPS thing is quite incredible from the ability to meet people!
I hung out with folks from Google, Meta, Microsoft, Apple, Weighs & Biases, Stability, Mistral, HuggingFace and PHD students and candidates from most of the top universities in the world, from KAIST to MIT and Stanford, Oslo and Shaghai, it's really a worldwide endeavor!
I also got to meet many of the leading figures in AI, all of whom I had to come up to and say hi, shake their hand, introduce myself (and ThursdAI) and chat about what they or their team released and presents at the conference! Truly an unforgettable experience!
Of course, This Weeks’ Buzz is that, everyone here loves W&B, from the PHD students, to literally every big LLM lab! They all came up to us (yes yes, even researches at Google who kinda low-key hate their internal tooling) and told us how awesome the experience was! (besides Xai folks, Jimmy wasn’t that impressed haha) and of course I got to practice the pitch so many times, since I manned the W&B booth!
Please do listen to the above podcast, there’s so much detail that’s in there that doesn’t get up on the newsletter, as it’s impossible to cover all, but it was a really fun conversation, including my excited depiction of this weeks NOLA escapades!
I think I’ll end here, cause I can go on and on about the parties (There were literally 7 at the same time last night, Google, Stability, OpenAI, Runway, and I’m sure there were a few more I wasn’t invited to!) and about New Orleans food (it’s my first time here, I ate a soft shell deep fried crab and turtle soup!) and I still have the poster sessions to go to and workshops! I will report more on my X account and the Weights & Biases X account, so stay tuned for that there, and as always, thanks for tuning in, reading and sharing ThursdAI with your friends 🫡
P.S - Still can’t really believe I get to do this full time now and share this journey with all of you, bringing you all with me to SF, and now NeurIPS and tons of other places and events in the future!
— Alex Volkov, AI Evangelist @ Weights & Biases, Host of ThursdAI 🫡
ThursdAI December 7th TL;DR
Greetins of the day everyone (as our panelist Akshay likes to sometimes say) and Happy first candle of Hannukah for those who celebrate! 🕎
I'm writing this newsletter from the back of an Waymo self driving car, in SF, as I'm here for just a few nights (again) to participate in the Open Source AI meetup, that was co-organized by Ollama and Nous Research, Alignment Labs and hosted by A16Z in their SF office.
This event was the highlight of this trip, it was quite a packed meetup in terms of AI talent, and I got to meet quite a few ThursdAI listeners, mutuals on X, and AI celebs
We also recorded the podcast this week from the arena, thanks to Swyx and Alessio from latentspace pod for hosting ThursdAI this week form their newly built out pod studio (and apologies everyone for the rocky start and the cutting out issues, luckily we had local recordings so the pod version sounds good!)
Google finally teases Gemini Ultra (and gives us Pro)
What a week folks, what a week, as I was boarding the flight to SF to meet with Open Source folks, Google announced (finally!) the release of Gemini, their long rumored, highly performant model with a LOT of fanfare!
Blogposts authored by Sundar and Demis Hassabis, beautiful demos of unseen before capabilities, comparisons to GPT-4V which the Ultra version of Gemini outperforms on several benchmarks, and rumors that Sergey Brin, the guy who's net worth is north of 100Bn is listed as the core contributor on the paper and reports on benchmarks (somewhat skewed) show Ultra beaing GPT-4 on many coding and reasoning evaluations!
We've been waiting for Gemini for such a long time, that we spend the first hour of the podcast discussing it and it's implications basically. We were also fairly disillusioned by the sleight of hand tricks Google marketing department played with the initial launch video, where it purportedly shows Gemini being a fully multi-modal AI, that reacts to a camera feed + user voice in real time, when in fact, it was quickly clear (from their developer blog) that it was not video+audio but rather images+text (the same two modalities we already have in GPT-4V and given some prompting, it's quite easy to replicate most of it. We've also discussed how we again, got a tease, and not even a waitlist for the "super cool" stuff, while getting a GPT3.5 level of a model today in Bard upgrade.
To me, the most mind-blowing demo video was actually one of the other ones in the announcement, which showed that Gemini has agentic behavior in understanding user intent, asks for clarifications, creates a PRD (Product Requirement Document) for itself, and then, generates Flutter code to create a UI on the fly, based on what the use asked it! This is pretty wild, as we all should expect that Just In Time UI will come to many of these big models!
Tune in to the episode if you want to hear more takes, opinions and frustrations as none of us actually got to use Gemini Ultra, and the experience with Gemini Pro (which is now live on Bard) was at least for me, underwhelming
This weeks buzz (What I learned in Weights & Biases this week)
I actually had a blast talking about W&B to many of the open source and fine-tuners community this and past week. I already learned that W&B doesn't only help huge companies (like OpenAI, Anthropic, Meta, Mistral and tons more) to train their foundational models, but is widely used by the open source fine-tuners community as well. I've met with folks like Wing Lian (aka Caseus), maintainer of Axolotl, who uses W&B together with Axolotl, and got to geek out about W&B, met with Teknium and LDJ (Nous Research, Alignment Labs) and in fact, got LDJ to walk me through some of the ways he uses and used W&B in the past, including how it's used to track model runs, show artifacts in the middle of runs, and run mini-benchmarks and evaluations for LLMS as they finetune.
If you're interested in this, here's an episode of a new “series” of me learning publicly (from scratch) so if you want to learn from scratch with me, welcome to check it out:
Open Source AI in SF meetup
This meetup was the reason I flew in to SF, I was invited by dear friends in the open source community, and couldn't miss it! There was such a talent density there, it was quite remarkable. Andrej Karpathy who's video about LLM I just finished re-watching, Jeremy Howard, folks from Mistral, A16Z, and tons of other startups, open source collectives, and enthusiasts, all came together to listen to a few lightning talks, but mostly to mingle and connect and share ideas.
Nous Research announced that they are a company (not anymore just a discord collective of rag tag open sourcers!) and that they are working on Forge, a product offering of theirs, that runs local AI, has a platform for agent behavior, and is very interesting to keep an eye for.
I've spent most of my time going around, hearing what folks are using (Hint: a LOT of axolotl), what they are finetuning (mostly Mistral) and what is the future (everyone's waiting for next Llama or next Mistral). Funnily enough, there was not a LOT of conversation about Gemini there at all, at least not among the folks that I talked to!
Overall this was really really fun, and of course, being in SF, at least for me, especially now as an AI Evangelist, feels like coming home! So expect more trip reports!
Here's a recap and a few more things that happened this week in AI:
* Open Source LLMs
* Apple released MLX - machine learning framework on apple silicon
* Mamba - transformers alternative architecture from Tri Dao
* Big CO LLMs + APIs
* Google Gemini beats GPT-4V on a BUNCH of metrics, shows cool fake multimodal demo
* Demo was embellished per the google developer blog
* Multimodal capabilities are real
* Dense model vs MOE
* Multimodel on the output as well
* For 5-shot, GPT-4 outperforms Gemini Ultra on MMLU
* AlphaCode 2 is here and Google claims it performs better than 85% competitive programmers in the world and it performs even better, collaborating with a competitive programmer.
* Long context prompting for Claude 2 shows 27% - 98% increase by using prompt techniques
* X.ai finally released grok to many premium+ X subscribers. (link)
* Vision
* OpenHermes Vision finally released - something there was not right there, back to drawing board
* Voice
* Apparently Gemini beats Whisper v3! As part of a unified model no less
* AI Art & Diffusion
* Meta - releases a standalong EMU AI art generator websites https://imagine.meta.com
* Tools
* Jetbrains finally releases their own AI native companion + subscription
That's it for me this week, this Waymo ride took extra long as it seems that in SF, during night rush hour, AI is at a disadvatage against human drivers. Maybe I'll take an Uber next time.
P.S - here’s Grok roasting ThursdAI
See you next week, and if you've scrolled all the way here for the emoji of the week, it's hidden in the middle of the article, send me that to let me know you read through 😉
🎶 Happy birthday to you, happy birthday to you, happy birthday chat GPT-eeeeeeee, happy birthday to you.
Hey everyone, welcome to this special edition of ThursdAI where you're probably gonna have two emails and two podcast episodes today and you can choose which one you want to but we actually recorded both of them live it just they went a little long.
ThursdAI - The only podcast that brings you yearly recaps since chatGPT was released (😂)
This one is the more celebratory one, today is one year from the release of chat GPT and we (and by we I mean I, Alex) decided to celebrate it by recapping not just the last week in AI but the last year (full timeline posted at the bottom of this newsletter)
Going month by month with a swoosh sound in the editing and covering the most important thing that happened in LLM and open source LLMs since chatGPT was released and imagination unlocked the capability for everyone!
We also covered Meta stepping in with Lama and then everything that happened since then in the multi modality and vector databases and agents and everything everything everything, it was a one hell of an hour and a half, we had almost 1K audience members! and so I recommend you listen to this one first and then the week updates later because there were some incredible releases this week as well! (as there are every week)
I think it's important to do like a Spotify wrapped type thing for AI, for something like a one year for chat GPT and I think we'll be doing this every year so hopefully in the year we'll see you here on November 30th covering the next year in AI.
And hopefully the next year in AI system will actually help me summarize all this because it's a lot of work but with that I will just leave you with the timeline and no notes and you should listen to everything because we talked about everything live!
I hope you enjoy this special birthday celebration! (OpenAI sure did, check out this incredibly cute little celebration video they just posted)
Here’s the full timeline with everything important that happened month by month that we’ve covered:
* December 2022 - ChatGPT becomes the fastest growing product in history
* GPT3.5 with 4K context window, instruction finetuning and conversational RLHF
* January
* Microsoft invests additional $10B into OpenAI (Jan 23, Blog)
* February
* LLaMa 1 - Biggest Open Source LLM (February 24 - Blog)
* No commercial license
* 30% MMLU
* No instruction fine-tuninig (RL;HF)
* ChatGPT unofficial APIs exist
* March (the month of LLM superpowers)
* ChatGPT API (March 1, announcement)
* Developers can now build chatGPT powered apps
* All clones so far were completion based and not conversation based
* LLama.cpp from ggerganov + Quantization (March 10, Blog)
* Stanford - Alpaca 7B - Finetune on self-instruct GPT3.5 dataset (March 13, Blog)
* GPT4 release + chatGPT upgrade (March 14 - GPT-4 demo)
* 67.0% HumanEval | 86.4% MMLU
* 8K (and 32K) context windows
* Anthropic announces Claude + Claude instant (March 14 - Blog)
* 56.0% HumanEval
* Folks previously form OAI leave to open Anthropic as research, then pivot from research into commercial
* LMSYS Vicuna 13B - Finetuned based on shareg.pt exports (March 30, Blog)
* April (Embedings & Agents)
* AutoGPT becomes the fastest github starred project + writes it's own code (April 1, Blog)
* Agents start to pop up like mushrooms after the rain
* LLaVa - Multimodality open source begins (April 18, Blog)
* CLIP + Vicuna smushed together to get LLMs eyes
* Bard improvements
* May (Context windows)
* Mosaic MPT-7B with 64K context, 1T parameters, commercial license (May 5, Blog)
* Anthropic updates Claude with 100K context window (May 11, Blog)
* LLongBoi summer begins (Context windows are being stretched)
* Nvidia shows Voyager agents that play Minecraft + Memory stored in Vector DB (May 27, Blog)
* June
* GPT-3.5-turbo + functions API (June 6, Blog)
* GPT3.5 and 4 got a boost in capabilities and steer-ability
* Price reduction on models + 75% reduction on ada embeddings model
* LLaMa context window extended to 8K with RoPE scaling
* AI Engineers self determination essay by swyx
* July
* Code Interpreter GA - ChatGPT can code (July 11, Blog)
* Anthropic Claude 2 - (July 11 - Blog)
* 200K context window
* 71% HumanEval
* LLaMa 2 (July 18 - Blog)
* Base & Chat models (RLHF)
* Commercial license
* 29.9% Human Eval | 68.9% MMLU
* August
* Meta releases Code-LlaMa, code finetune models
* September
* DALL-E 3 - Adds multi-modality on output and chat to image gen (Sep 20, Blog)
* Mistral 7B top performing open source LLM via torrent link (Sep 27, Blog)
* GPT4-V (vision & voice) - Adds multimodality on input (Sep 27, Blog)
* October
* OpenHermes - Mistral 7B finetune that tops the charts from Teknium / Nous Research (Oct 16, Announcement)
* Inflection PI gets connected to the web + supportPi mode (Oct 16, Blog)
* Adept releases multimodal FuYu 8B (Oct 19, blog)
* November
* Grok from Xai - with realtime access to all of X content
* OpenAI dev day
* Combined mode for MMIO (multi modal on input and output)
* GPT-4 Turbo with 128K context, 3x cheaper than GPT-4
* Assistants API with retrieval capabilities
* Share-able GPTs - custom versions of GPT with retrieval, DALL-E, Code Interpreter and vision
* Chatbots with real business use-cases, for example WandBot (that we just launched today! Blog)
* Has vector storage memory
* Available via Discord/Slack
* And custom GPT!
* Microsoft has copilot everywhere in office
Aaaand now we’re here!
What an incredible year, can’t imagine what the next year holds for all of us, but 1 thing is for sure, ThursdAI will be here to keep you all up to date!
P.S - If you scrolled all the way to here, DM me the 🎊 emoji so I know you celebrated with us! It really helps me to know that there is at least a few folks out of the thousands that get this newsletter that scrolls all the way through!
ThursdAI TL;DR - November 23
TL;DR of all topics covered:
* OpenAI Drama
* Sam... there and back again.
* Open Source LLMs
* Intel finetuned Mistral and is on top of leaderboards with neural-chat-7B (Thread, HF, Github)
* And trained on new Habana hardware!
* Yi-34B Chat - 4-bit and 8-bit chat finetune for Yi-34 (Card, Demo)
* Microsoft released Orca 2 - it's underwhelming (Thread from Eric, HF, Blog)
* System2Attention - Uses LLM reasons to figure out what to attend to (Thread, Paper)
* Lookahead decoding to speed up LLM inference by 2x (Lmsys blog, Github)
* Big CO LLMs + APIs
* Anthropic Claude 2.1 - 200K context, 2x less hallucinations, tool use finetune (Announcement, Blog, Ctx length analysis)
* InflectionAI releases Inflection 2 (Announcement, Blog)
* Bard can summarize youtube videos now
* Vision
* Video-LLaVa - open source video understanding (Github, demo)
* Voice
* OpenAI added voice for free accounts (Announcement)
* 11Labs released speech to speech including intonations (Announcement, Demo)
* Whisper.cpp - with OpenAI like drop in replacement API server (Announcement)
* AI Art & Diffusion
* Stable Video Diffusion - Stability releases text2video and img2video (Announcement, Try it)
* Zip-Lora - combine diffusion LORAs together - Nataniel Ruiz (Annoucement, Blog)
* Some folks are getting NERFs out from SVD (Stable Video Diffusion) (link)
* LCM everywhere - In Krea, In Tl;Draw, in Fal, on Hugging Face
* Tools
* Screenshot-to-html (Thread, Github)
Ctrl+Altman+Delete weekend
If you're subscribed to ThursdAI, then you most likely either know the full story of the crazy OpenAI weekend. Here's my super super quick summary (and if you want a full blow-by-blow coverage, Ben Tossel as a great one here)
Sam got fired, Greg quit, Mira flipped then Ilya Flipped. Satya played some chess, there was an interim CEO for 54 hours, all employees sent hearts then signed a letter, neither of the 3 co-fouders are on the board anymore, Ilya's still there, company is aligned AF going into 24 and Satya is somehow a winner in all this.
The biggest winner to me is open source folks, who got tons of interest suddenly, and specifically, everyone seems to converge on the OpenHermes 2.5 Mistral from Teknium (Nous Research) as the best model around!
However, I want to shoutout the incredible cohesion that came out of the folks in OpenAI, I created a list of around 120 employees on X and all of them were basically aligned the whole weekend, from ❤️ sending to signing the letter, to showing how happy they are Sam and Greg are back!
Yay
This Week's Buzz from WandB (aka what I learned this week)
As I’m still onboarding, the main things I’ve learned this week, is how transparent Weights & Biases is internally. During the whole OAI saga, Lukas the co-founder sent a long message in Slack, addressing the situation (after all, OpenAI is a big customer for W&B, GPT-4 was trained on W&B end to end) and answering questions about how this situation can affect us and the business.
Additionally, another co-founder, Shawn Lewis shared a recording of his update to the BOD of WandB, about out progress on the product side. It’s really really refreshing to see this information voluntarily shared with the company 👏
The first core value of W&B is Honesty, and it includes transparency outside of matters like personal HR stuff, and after hearing about this during onboarding, it’s great to see that the company lives it in practice 👏
I also learned that almost every loss curve image that you see on X, is a W&B dashboard screenshot ✨ and while we do have a share functionality, it’s not built for viral X sharing haha so in the spirit of transparency, here’s a video I recorded and shared with product + feature request to make these screenshot way more attractive + clear that it’s W&B
Open Source LLMs
Intel passes Hermes on SOTA with a DPO Mistral Finetune (Thread, Hugging Face, Github)
Yes, that intel, the... oldest computing company in the world, not only comes out strong with the best (on benchmarks) open source LLM, it also does DPO, and has been trained on a completely new hardware + Apache 2 license!
Here's Yam's TL;DR for the DPO (Direct Policy Optimization) technique:
Given a prompt and a pair of completions, train the model to prefer one over the other. This model was trained on prompts from SlimOrca's dataset where each has one GPT-4 completion and one LLaMA-13B completion. The model trained to prefer GPT-4 over LLaMA-13B.
Additionally, even tho there is custom hardware included here, Intel supports the HuggingFace trainer fully, and the whole repo is very clean and easy to understand, replicate and build things on top of (like LORA)
LMSys Lookahead decoding (Lmsys, Github)
This method significantly improves the output of LLMs, sometimes by more than 2x, using some jacobian notation (don't ask me) tricks. It's copmatible with HF transformers library! I hope this comes to open source tools like LLaMa.cpp soon!
Big CO LLMs + APIs
Anthropic Claude comes back with 2.1 featuring 200K context window, tool use
While folks on X thought this was new, Anthropic actually announced Claude with 200K back in the May, and just gave us 100K context window, which for the longest time was the largest context window around. I was always thinking, they don't have a reason to release 200K since none of their users actually wants it, and it's a marketing/sales decision to wait until OpenAI catches up. Remember, back then, GPT-4 was 8K and some lucky folks got 32K!
Well, OpenAI releases GPT-4-turbo with 128K so Anthropic re-trained and released Claude to gain an upper hand. I also love the tool use capabilities.
Re: longer context window, there were a bunch of folks testing if 200K context window is actually all that great, and it turns out, besides being very expensive to run (you pay per tokens) it also loses a bunch of information at lengths over 200K using needle in the haystack searches. Here's an analysis by Greg Kamradt that shows that:
* Starting at ~90K tokens, performance of recall at the bottom of the document started to get increasingly worse
* Less context = more accuracy - This is well know, but when possible reduce the amount of context you send to the models to increase its ability to recall.
I had similar issues back in May with their 100K tokens window (source)
Voice & Audio
ElevenLabs has speech-to-speech
Creating a significant jump in capabilities, ElevenLabs now allows you to be the actor behind the voice! With speech to speech, they would transfer the pauses, the intonation, the emotion, into the voice generation. Here's my live reaction and comparison:
* Notable: Whisper.CPP now supports a server compatible with OpenAI (Announcement, Github)
AI Art & Diffusion
Stable Video diffusion - text-2-video / img-2-video foundational model (Announcement, Hugging Face, Github, DEMO)
Stable has done it again, Stable Video allows you to create increidbly consistent videos with images or just text! They are short now, but working on extending the times, and they videos look incredible! (And thanks to friends at Fal, you can try right now, here)
And here’s a quick gif I created with DALL-E 3 and Fal to celebrate the Laundry Buddy team at OAI while the outage was happening)
Tools
Screenshot to HTML (Github)
I… what else is there to say? Someone used GPT4-Vision to … take screenshots and iteratively re-create the HTML for them. As someone who used to spend month on this exact task, I’m very very happy it’s now automated!
Happy Thanksgiving 🦃
I am really thankful to all of you who subscribe and come back every week, thank you! I would have been here without all your support, comments, feedback! Including this incredible art piece that Andrew from spacesdashboard created just in time for our live recording, just look at those little robots! 😍
See you next week (and of course the emoji of the week is 🦃, DM or reply!)
Hey yall, welcome to this special edition of ThursdAI. This is the first one that I'm sending in my new capacity as the AI Evangelist Weights & Biases (on the growth team)
I made the announcement last week, but this week is my first official week at W&B, and oh boy... how humbled and excited I was to receive all the inspiring and supporting feedback from the community, friends, colleagues and family 🙇♂️
I promise to continue my mission of delivering AI news, positivity and excitement, and to be that one place where we stay up to date so you don't have to.
This week we also had one of our biggest live recordings yet, with 900 folks tuned in so far 😮 and it was my pleasure to again to chat with folks who "made the news" so we had a brief interview with Steve Ruiz and Lou from TLDraw, about their incredible GPT-4 Vision enabled "make real" functionality and finally got to catch up with my good friend Idan Gazit who's heading the Github@Next team (the birthplace of Github Copilot) about how they see the future. So definitely definitely check out the full conversation!
TL;DR of all topics covered:
* Open Source LLMs
* Nous Capybara 34B on top of Yi-34B (with 200K context length!) (Eval, HF)
* Microsoft - Phi 2 will be open sourced (barely) (Announcement, Model)
* HF adds finetune chain genealogy (Announcement)
* Big CO LLMs + APIs
* Microsoft - Everything is CoPilot (Summary, copilot.microsoft.com)
* CoPilot for work and 365 (Blogpost)
* CoPilot studio - low code "tools" builder for CoPilot + GPTs access (Thread)
* OpenAI Assistants API cookbook (Link)
* Vision
* 🔥 TLdraw make real button - turn sketches into code in seconds with vision (Video, makereal.tldraw.com)
* Humane Pin - Orders are out, shipping early 2024, multimodal AI agent on your lapel (
* )
* Voice & Audio
* 🔥 DeepMind (Youtube) - Lyria high quality music generations you can HUM into (Announcement)
* EmotiVoice - 2000 different voices with emotional synthesis (Github)
* Whisper V3 is top of the charts again (Announcement, Leaderboard, Github)
* AI Art & Diffusion
* 🔥 Real-time LCM (latent consistency model) AI art is blowing up (Krea, Fal Demo)
* 🔥 Meta announces EMU-video and EMU-edit (Thread, Blog)
* Runway motion brush (Announcement)
* Agents
* Alex's Visual Weather GPT (Announcement, Demo)
* AutoGen, Microsoft agents framework is now supporting assistants API (Announcement)
* Tools
* Gobble Bot - scrape everything into 1 long file for GPT consumption (Announcement, Link)
* ReTool state of AI 2023 - https://retool.com/reports/state-of-ai-2023
* Notion Q&A AI - search through a company Notion and QA things (announcement)
* GPTs shortlinks + analytics from Steven Tey (
https://chatg.pt
* )
This Week's Buzz from WandB (aka what I learned this week)
Introducing a new section in the newsletter called "The Week's Buzz from WandB" (AKA What I Learned This Week).
As someone who joined Weights and Biases without prior knowledge of the product, I'll be learning a lot. I'll also share my knowledge here, so you can learn alongside me. Here's what I learned this week:
The most important things I learned this week is just how prevelant and how much of a leader Weights&Biases is. W&B main product is used by most of the foundation LLM trainers including OpenAI.
In fact GPT-4 was completely trained on W&B!
It's used by pretty much everyone besides Google. In addition to that it's not only about LLMs, W&B products are used to train models in many many different areas of the industry.
Some incredible examples are a pesticide dispenser that's part of the John Deere tractors that only spreads pesticides onto weeds and not actual produce. And Big Pharma who's using W&B to help create better drugs that are now in trial. And it's just incredible how much machine learning that's outside of just LLMs is there. But also I'm absolutely floored by just the amount of ubiquity that W&B has in the LLM World.
W&B has two main products, Models & Prompts, Prompts is a newer one, and we're going to dig into both of these more next week!
Additionally, it's striking how many AI Engineers, API users such as myself and many of my friends, have no idea of who W&B even is, of if they do, they never used it!
Well, that's what I'm here to change, so stay tuned!
Open source & LLMs
In the open source corner, we have the first Nous fine-tune of Yi-34B, which is a great model that we've covered in the last episode and now is fine-tuned with the Capybara dataset by ThursdAI cohost, LDJ! Not only is that a great model, it now tops the charts for the resident reviewer we WolframRavenwolf on /r/LocalLLama (and X)
Additionally, Open-Hermes 2.5 7B from Teknium is now second place on HuggingFace leaderboards, it was released recently but we haven't covered until now, I still think that Hermes is one of the more capable local models you can get!
Also in open source this week, guess who loves it? Satya (and Microsoft)
They love it so much that they not only created this awesome slide (altho, what's SLMs? Small Language Models? I don't like it), they also announced that LLaMa and Mistral are coming to Azure services as inference!
And they gave us a little treat, Phi2 is coming. They said OpenSource (but folks looking at the license saw that it's only for research capabilities) but supposedly it's a significantly more capable model while only being 2.7B weights (super super tiny)
Big Companies & APIs
Speaking of Microsoft, they announced so much during their Ignite event on Wednesday (15th) that it's impossible to cover all of it in this newsletter, but basically here are the main things that got me excited!
CoPilot everywhere, everything is CoPilot
Microsoft rebranded Bing Chat to Copilot and it now lives on copilot.microsoft.com
and it's basically a free GPT4, with vision and DALL-e capabilities. If you're not signed up for OpenAI's plus membership, this is as good as it gets for free!
They also announced CoPilot for 365, which means that everything from office (word, excel!) to your mail, and your teams conversations will have a CoPilot that will be able to draw from your organizations knowledge and help you do incredible things. Things like help book appointments, pull in relevant people for the meeting based on previous documents, summarize that meeting, schedule follow ups, and like a TON more stuff. Dall-e integration will help you create awesome powerpoint slides.
(p.s. all of this will be allegedly data protected and won't be shared with MS or be trained on)
They literally went and did "AI everywhere" with CoPilot and it's kinda incredible to see how big they are betting the farm on AI with Microsoft while Google... where's Google™?
CoPilot Studio
One of the more exciting things for me was, the CoPilot Studio announcement, a low-code tool to extend your company's CoPilot by your IT, for your organization. Think, getting HR data from your HR system, or your sales data from your SalesForce!
They will launch with 1100 connectors for many services but allow you to easily build your own.
One notable thing is, Custom GPTs will also be a connector! You will be literally able to connect yyour CoPilot with your (or someone's) GPTs! Are you getting this? AI Employees are coming faster than you think!
Vision
I've been waiting for cool vision demos since GPT-4V API was launched and oh boy did we get them! From friend of the pod Robert Lukoshko Auto screenshot analysis which will take screenshots periodically and will send you a report of all you did that day, to Charlie Holtz live webcam narration by David Attenborough (which is available on Github!)
But I think there's 1 vision demo that takes the cake this week, by our friends (Steve Ruiz) from TLDraw, which is a whiteboard canvas primitive. They have added a sketch-to-code button, that allows you to sketch something out and GPT-4 Vision will analyze this, and GPT-4 will write code, and you will get live code within seconds. It's so mind-blowing that I'm still collecting my jaw of the floor here. It also does coding, so if you ask it nicely to add JS interactivity, the result will be interactive 🤯
GPT4-V Is truly as revolutionary as I imagined it to be when Greg announced it on stage 🫡
P.S - Have you played with it? Do you have cool demos? DM me with 👁️🗨️ emoji and a cool vision demo to be included in the next ThursdAI
AI Art & Diffusion & 3D
In addition to the TLDraw demo, one mind-blowing demo after another is coming this week from the AI Art world, using the LCM (Latent Consistency Model) + a whiteboard. This is yet another see it to believe it type thing (or play with it)
(video from Linus)
Dear friends from Krea.ai were the first to implement this insanity, that allows you to see real time AI art generation almost as fast as you type your prompts, and then followed up by the wizards at Fal to get the generations down to several mili-seconds (shoutout Gorkem!), the real time drawing thing is truly truly mind-blowing. It's so mind-blowing that folks add their webcam feeds into this, and see almost real time generation on the fly of their webcam feeds.
Meta announcing new Emus (Video & Edit)
Meta doesn't want me to relax, and during the space, announced their text-to-video and textual-editing models.
Emu Video produces great videos from a prompt, and emu-edit is really interesting, it allows you to edit parts of images by typing, think "remove the tail from this cat" or "remove the hat from this person"
They have this to say, which... dayum.
In human evaluations, our generated videos are strongly preferred in quality compared to all prior work– 81% vs. Google’s Imagen Video, 90% vs. Nvidia’s PYOCO, and 96% vs. Meta’s Make-A-Video. Our model outperforms commercial solutions such as RunwayML’s Gen2 and Pika Labs
It's really compelling, can't wait to see if they open source this, video is coming ya'll!
Audio & Sound
Deepmind + Youtube announced Lyria (blogpost)
This new music model is pretty breathtaking, but we only got a glimpse, not even a waitlist for that one, however, check out the pre-recorded demoes, folks at deep mind have a model you can hum into, sing into, and it'll create a full blown track for you, with bass, drums, and singing!
Not only that, it will also license vocals from mucisians (al-la Grimes) and will split the revenue between you and them if you post it on Youtube!
Pretty cool Google, pretty cool!
Agents & Tools
Look, I gotta be honest, I'm not sure about this category, Agent and Tools, if to put them into one or not, but I guess GPTs are kinda tools, so I'm gonna combine them for this one.
GPTs (My Visual Weather, Simons Notes)
This week, the GPT that I created called Visual Weather GPT has blown up, with over 5,000 chats opened with it, and many many folks using this and texting me about this. super cool way to just like check all the capabilities of a GPT. If you remember, I thought of this idea a few weeks ago when we got a sneak preview to the "All tools" mode, but now I can share it with you all in the form of a GPT, that will browser the web for real time weather data, and create a unique art piece for that location and weather conditions!
It's really easy to make as well, and I do fully expect everyone to start making their own versions very soon, and I think we're inching towards the era of JIT (just in time) software, where you'll create software as you require it, and it'll be as easy as talking to a chatGPT!
Speaking of, friend of the pod Steven Tey from Vercel (who's dub.sh I use and love for thursdai.news links) has released a GPT link shortener, called chatg.pt and you can register and get your own cool short link like https://chatg.pt/artweather 👏 And it'll give you analytics as well!
Pro tip for weather GPT, you can ask for a specific season or style in parentesises and then those as greeting cards for your friends. Happy upcoming Thanks giving everyone!
Speaking of Thanks
giving, we're not taking a break, next ThursdAI, November 23, join us for a live discussion and podcast recoding! We'll have many thanks, cool AI stuff, and much more!
Hey everyone, this is Alex Volkov 👋
This week was an incredibly packed with news, started strong on Sunday with x.ai GrŌk announcement, Monday with all the releases during OpenAI Dev Day, then topped of with Github Universe Copilot announcements, and to top it all of, we postponed the live recording to see what hu.ma.ne has in store for us as AI devices go (Finally announced Pin with all the features)
In between we had a new AI Unicorn from HongKong called Yi from 01.ai which dropped a new SOTA 34B model with a whopping 200K context window and a commercial license by ex-Google China lead Kai Fu Lee.
Above all, this week was a monumental for me personally, ThursdAI has been a passion project for the longest time (240 days), and it led me to incredible places, like being invited to ai.engineer summit to do media, then getting invited to OpenAI Dev Day (to also do podcasting from there), interview and befriend folks from HuggingFace, Github, Adobe, Google, OpenAI and of course open source friends like Nous Research, Alignment Labs, and interview authors of papers, hackers of projects, and fine-tuners and of course all of you, who tune in from week to week 🙏 Thank you!
It's all been so humbling and fun, which makes me ever more excited to share the next chapter. Starting Monday I'm joining Weights & Biases as an AI Evangelist! 🎊
I couldn't be more excited to continue ThursdAI mission, of spreading knowledge about AI, connecting between the AI engineers and the fine-tuners, the Data Scientists and the GEN AI folks, the super advanced cutting edge stuff, and the folks who fear AI with the backing of such an incredible and important company in the AI space.
ThursdAI will continue as a X space, newsletter and podcast, as we'll gradually find a common voice, and continue bringing folks awareness of WandB incredible brand to newer developers, products and communities. Expect more on this very soon!
Ok now to the actual AI news 😅
TL;DR of all topics covered:
* OpenAI Dev Day
* GPT-4 Turbo with 128K context, 3x cheaper than GPT-4
* Assistant API - OpenAI's new Agent API, with retrieval memory, code interpreter, function calling, JSON mode
* GPTs - Shareable, configurable GPT agents with memory, code interpreter, DALL-E, Browsing, custom instructions and actions
* Privacy Shield - Open AI lawyers will protect you from copyright lawsuits
* Dev Day emergency pod with Latent Space with Swyx, Allesio, Simon and Me! (Listen)
* OpenSource LLMs
* 01 launches YI-34B, a 200K context window model commercially licensed and it tops all HuggingFace leaderboards across all sizes (Announcement)
* Vision
* GPT-4 Vision API finally announced, rejoice, it's as incredible as we've imagined it to be
* Voice
* Open AI TTS models with 6 very-realistic, multilingual voices, no cloning tho
* AI Art & Diffusion
* <2.5 seconds full SDXL inference with FAL (Announcement)
OpenAI Dev Day
So much to cover from OpenAI that this has it's own section today in the newsletter.
I was lucky enough to get invited, and attend the first ever OpenAI developer conference (AKA Dev Day) and it was an absolute blast to attend. It was also incredible to attend it together with all 8.5 thousand of you who tuned into our live stream on X, as we were walking to the event, and then watched the keynote together (Thanks Ray for the restream) and talked with OpenAI folks about the updates. Huge shoutout to LDJ, Nisten, Ray, Phlo, Swyx and many other folks who held the space, while we were otherwise engaged with deep dives and meeting folks and doing interviews!
So now for some actual reporting! What did we get from OpenAI? omg we got so much, as developers, as users (and as attendees, I will add more on this later)
GPT4-Turbo with 128K context length
The major thing that was announced is a new model, GPT-4-turbo, which is supposedly faster than GPT-4, while being 3x cheaper (2x on output) and having a whopping 128K context length while also being more accurate (with significantly better recall and attention throughout this context length)
With JSON mode and significantly improved function calling capabilities, updated cut-off time (April 2023), and higher rate limits, this new model is already being implemented across all the products and is a significant significant upgrade to many folks
GPTs - A massive shift in agent landscapes by OpenAI
Another (semi-separate) thing that Sam talked about was the GPTs, their version of agents
not to be confused with the Assistants API, which is also Agents, but for developers, and they are not the same and it's confusing
GPTs I think is a genius marketing move by OpenAI and replaces Plugins (that didn't even meet product market fit) in many regards.
GPTs are instances of well... GPT4-turbo, that you can create by simply chatting with BuilderGPT, and they can have their own custom instruction set, and capabilities that you can turn on and off, like browse the web with Bing, Create images with DALL-E and write and execute code with Code Interpreter (bye bye Advanced Data Analysis, we don't miss ya).
GPTs also have memory, you can upload a bunch of documents (and your users as well) and GPTs will do vectorization and extract the relevant information out of those documents, so think, your personal Tax assistant that has all 3 years of your tax returns
And they have eyes, GPT4-V is built in so you can drop in screenshots, images and all kinds of combinations of things.
Additionally, you can define actions for assistants (which is similar to how Plugins were developed previously, via an OpenAPI schema) and the GPT will be able to use those actions to do tasks outside of the GPT context, like send emails, check stuff in your documentation and much more, pretty much anything that's possible via API is now possible via the actions.
One big thing that's missing for me is, GPTs are reactive, so they won't reach out to you or your user when there's a new thing, like a new email to summarize or a new task completed, but I'm sure OpenAI will close that gap at some point.
GPTs are not Assistants, they are similar but not the same and it's quite confusing. GPTs are created online, and then are share-able with links.
Which btw, I created a GPT that uses several of the available tools, browsing for real time weather info, and date/time and generates an on the fly, never seen before weather art for everyone. It's really fun to play with, let me know what you think (HERE) the image above is generated by the Visual Weather GPT
Unified "All tools" mode for everyone (who pays)
One tiny thing that Sam mentioned on stage, is in fact huge IMO, is the removal of the selector in chatGPT, and all premium users now have access to 1 interface that is multi modal on input and output (I call it MMIO) - This mode now understands images (vision) + text on input, and can browse the web and generate images, text, graphs (as it runs code) on the output.
This is a significant capabilities upgrade to many folks who will use these tools, but previously didn't want to choose between DALL-E mode and browse or Code Interpreter mode. The model now intelligently selects what tool to use for a given task, and this means more and more "generality" for the models, as they are learning and getting new capabilities in the form of tools.
This in addition to a MASSIVE 128K context window, means that chatGPT has now been significantly upgraded, and you still pay $20/mo 👏 Gotta love that!
Assistant API (OpenAI Agents)
This is the big announcement for developers, we all got access to a new and significantly improved Assistants API, which improves on several our experience on several categories:
* Creating Assistants - Assistants are OpenAI's first foray into the world of AGENTS, and it's quite exciting! You can create an assistant via API (not quite the same as GPTs, we'll cover the differences later), you can create each assistant with it's own set of instructions (that you don't have to pass each time with the prompt), tools like code interpreter and retrieval, and functions. Also you can select models, so you don't have to use the new GPT-4-turbo (but you should!)
* Code Interpreter - Assistants are able to write and execute code now, which is a whole world of excitement! Having code abilities (that executes in a safe environment on OAI side) is a significant boost in many regards and many tasks require bits of code "on the fly", for example time-zone tasks. You will no longer have to write that code yourself, you can ask your assistant
* Retrieval - OpenAI (and apparently QDrant!) have given all the developers a built in RAG (retrieval augmented generation) capabilities + document uploading and understanding. You can upload files like documentation via the API or let your users upload files, and parse and extract information out of! This is an additional huge huge thing, basically memory is built in for you now
* Stateful API - this API introduces the concept of threads, where OpenAI will manage the state of your conversation, and you can assign 1 user per thread and then just send the responses back to the user, and send the user queries to the same thread. No longer do you have to send the whole history back and forth! It's quite incredible, however it raises the question of pricing, and calculating tokens. Per OpenAI (I asked), if you would like to calculate costs on the fly, you'd have to use the get thread endpoint, and then count the amount of tokens that's already in the thread (and it can be a LOT since there's now 128K tokens in the context length)
* JSON and Better functions calling - You can now set the API to respond in JSON mode! Which is an incredible improvement for devs, and which we only were able to do via Functions before, and even functions got an upgrade, with an ability to call multiple functions. Functions are added as "actions" in the assistant creation, so you can give your assistant abilities that it will execute by returing to you functions with the right parameters. Thing "set the mood" will return a function to call the smart lights, and "play" will return a function that will call Spotify API
* Multiple Assistants can join a thread - you can create specific assistants that can all join the same thread with the user, each with a set of custom instructions and capabilities and tools
* Parallel Functions - this is also new, the assistant API can now return several functions for you to execute, which could lead to the creation of scenes, for example in a smart home, you want to "set the mood" and several functions would be returned from the API, one that will turn of the lights, one that will start the music, and one that will turn on mood lighting.
Vision
GPT-4 Vision
Finally, it's here, multimodality for developers to implement, the moment I personally have been waiting for since GPT-4 was launched (and ThursdAI started) back on March 14 (240 days ago, but who's counting)
GPT-4 vision is able to take images, and text, and respond with many vision related tasks, like analysis, understanding, summarization of captions. Many folks are splitting videos frame by frame and analyzing whole videos already (in addition to whispering the video to get what is said)
Hackers and developers like friend of the pod Robert, created quick hacks like a browser extension that lets you select any screenshot on the page and ask GPT4 vision things about it, another friend of the pod SkalskiP created a hot dog classifier Gradio space 😂 and is maintaining an awesome list of experiments with vision on Github
Voice
Text to speech models
OpenAI decided to help us all build agents properly, and agents need not only ears (for which they gave us whisper, and released V3 as well) but also voice, and we finally got the TTS from OpenAI, 6 very beautiful, emotional voices, that you can run very easily, and cheaply. You can't generate more or clone yet (that's only for friends of OpenAI like Spotify and others) but you can use the 6 we got (plus a secret pirate one apparently they trained but never released!)
They sound ultra-realistic, and are multi-linugal as well, you can just pass different languages and voila. Friend of the pod Simon Willison created a quick CLI tool called ospeak to pipe text into and it'll use your OAI key to read that text out with those super nice voices!
Whisper v3 was released!
https://github.com/openai/whisper/discussions/1762
The large-v3 model shows improved performance over a wide variety of languages, and the plot below includes all languages where Whisper large-v3 performs lower than 60% error rate on Common Voice 15 and Fleurs, showing 10% to 20% reduction of errors compared to large-v2:
HUMANE
Humane AI pin is ready for pre-order at 699
HUMANE pin was finally announced, and here is the break-down, they have a clever way to achieve "all day battery life" by having a hot swap system, with a magnetic booster that you can swap when you get low on battery (pretty genius TBH)
It's passive so it's not "always listening" but there is a wake word apparently, and you can activate by touch. Runs on the T-mobile Network ( which sucks for folks like me where T-mobile just doesn't have reception in their neighborhood 😂 )
No apps, just AI experiences powered by OpenAI, with a laser powered projector UI on your hand, and voice controls
AI voice input will allow interactions like asking for information (which has browsing) and is SIGNIFICANTLY better than "Siri" or "Ok Google" from the demo, being able to rewrite your messages for you, catch you up on multiple messages and even search through them! You can ask for retrieval from previous messages
Pin is multimodal, voice input and vision
Holding the microphone on Tab while someone's speaking to you in a different language will automatically translate that language for you and then translate you back to that language with your own intonation! Bye bye language barriers!
And with vision, you can do things like tracking calories from showing it what you ate, or buy things you're seeing in the store, but online, take pictures and videos and then store all of them transcribed in your personal AI memory
Starting at $699, with a $24/mo payment that comes with unlimited AI queries, storage and service (again, just T-mobile), Tidal music subscription and more.
I think it's lovely that someone tries to take on Google/Apple duopoly with a completely re-imagined AI device, and can't wait to pre-order mine and test it out. It will be an interesting experience of balance with 2 phone numbers, but also a monthly payment that basically makes the device use-less if you stop paying.
Phew, this was a big update, not to mention there's a whole 2 hour podcast I want you to listen to on top of this, thank you for reading, for subscribing, for participating in the community and I can't wait to finally relax after this long week (still Jet-lagged) and prepare for my new Monday!
I want to send a heartfelt shoutout to my friend swyx who not only let me on to Latent Space from time to time (including the last recap emergency pod), but also is my blood-line to SF, where everything happens! Thanks man, I really appreciate all you did for me and ThursdAI 🫡
Can't wait to see you all on the next ThursdAI, and as always, replies, comments, congratulations, are welcome as replies, DMs and send me the 🎉 for this one, I'd really appreciate it!
— Alex
ThursdAI November 2nd
Hey everyone, welcome to yet another exciting ThursdAI. This week we have a special announcement, the co-host of and I will be hosting a shared X space live from Open AI Dev Day! Monday next week (and then will likely follow up with interviews, analysis and potentially a shared episode!)
Make sure you set a reminder on X (https://thursdai.news/next) , we’re going to open the live stream early, 8:30am on Monday, and we’ll live stream all throughout the keynote! It’ll be super fun!
Back to our regular schedule, we covered a LOT of stuff today, and again, were lucky enough to have BREAKING NEWS and the authors of said breaking news (VB from HuggingFace and Emozilla from Yarn-Mistral-128K) to join us and talk a little bit in depth about their updates!
[00:00:34] Recap of Previous Week's Topics
[00:00:50] Discussion on AI Embeddings
[00:01:49] Gradio Interface and its Applications
[00:02:56] Gradio UI Hosting and its Advantages
[00:04:50] Introduction of Baklava Model
[00:05:11] Zenova's Input on Distilled Whisper
[00:10:32] AI Regulation Week Discussion
[00:24:14] ChatGPT new All Tools mode (aka MMIO)
[00:35:45] Discussion on Multimodal Input and Output Models
[00:36:55] BREAKING NEWS: Mistral YaRN 7B - 128K context window
[00:37:02] Announcement of Mistral Yarn Release
[00:46:47] Exploring the Limitations of Current AI Models
[00:47:25] The Potential of Vicuna 16k and Memory Usage
[00:49:43] The Impact of Apple's New Silicon on AI Models
[00:51:23] Introduction to New Models from Nius Research
[00:51:39] The Future of Long Context Inference
[00:53:42] Exploring the Capabilities of Obsidian
[00:54:29] The Future of Multimodality in AI
[00:58:48] The Exciting Developments in CodeFusion
[01:06:49] The Release of the Red Pajama V2 Dataset
[01:12:07] The Introduction of Luma's Genie
[01:16:37] Discussion on 3D Models and Stable Diffusion
[01:17:08] Excitement about AI Art and Diffusion Models
[01:17:48] Regulation of AI and OpenAI Developments
[01:18:24] Guest Introduction: VB from Hug& Face
[01:18:53] VB's Presentation on Distilled Whisper
[01:21:54] Discussion on Distillation Concept
[01:27:35] Insanely Fast Whisper Framework
[01:32:32] Conclusion and Recap
Show notes and links:
* AI Regulation
* Biden Executive Order on AI was signed (Full EO, Deep dive)
* UK AI regulation forum (King AI speech, no really, Arthur from Mistral)
* Mozilla - Joint statement on AI and openness (Sign the letter)
* Open Source LLMs
* Together AI releases RedPajama 2, 25x larger dataset (30T tokens) (Blog, X, HF)
* Alignment Lab - OpenChat-3.5 a chatGPT beating open source model (HF)
* Emozilla + Nous Research - Yarn-Mistral-7b-128k (and 64K) longest context window (Announcement, HF)
* LDJ + Nous Research release Capybara 3B & 7B (Announcement, HF)
* LDJ - Obsidian 3B - the smallest open source multi modal model (HF, Quantized)
* Big CO LLMs + APIs
* ChatGPT "all tools" MMIO mode - Combines vision, browsing, ADA and DALL-E into 1 model (Thread, Examples, System prompt)
* Microsoft CodeFusion paper - a tiny (75M parameters) model beats a 20B GPT-3.5-turbo (Thread, ArXiv)
* Voice
* Hugging Face - Distill whisper - 2x smaller english only version of Whisper (X, paper, code)
* AI Art & Diffusion & 3D
* Luma - text-to-3D Genie bot (Announcement, Try it)
* Stable 3D & Sky changer
AI Regulation IS HERE
Look, to be very frank, I want to focus ThursdAI on all the news that we're getting from week to week, and to bring a positive outlook, so politics, doomerism, and regulation weren't on the roadmap, however, with weeks like these, it's really hard to ignore, so let's talk about this.
President Biden signed an Executive Order, citing the old, wartime era Defence Production act (looks like the US gov. also has "one weird trick" to make the gov move faster) and it wasn't as bombastic as people thought. X being X, there has been so many takes pre this executive order even releasing about regulatory capture being done by the big AI labs, about how open source is no longer going to be possible, and if you visit Mark Andressen feed you'll see he's only reposting AI generated memes to the tune of "don't tread on me" about GPU and compute rights.
However, at least on the face of it, this executive order was mild, and discussed many AI risks and focused on regulating models from huge compute runs (~28M H100 hours // $50M dollars worth). Here's the relevant section.
Many in the open source community reacted to the flops limitation with a response that it's very much a lobbyist based decision, and that the application should be regulated, not only the compute.
There's much more to say about the EO, if you want to dig deeper, I strongly recommend this piece from AI Snake oil :
and check out Yan Lecun's whole feed.
UK AI safety summit in Bletchley Park
Look, did I ever expect to add the King of England into an AI weekly recap newsletter? Surely, if he was AI Art generated or something, not the real king, addressing the topic of AI safety!
This video was played for the attendees of a few day AI safety summit in Blecheley park, where AI luminaries (Yan Lecun, Elon Musk, Arthur Mensch Mistral CEO, Naveen Rao) attended and talked about the risks and benefits of AI and regulation. I think Naveen Rao had a great recap here, but additionally, there were announcements about Safety Institute in the UK, and they outlined what actions the government can take.
In other regulation related news, Mozilla has a joint statement on AI safety and openness (link) that many signed, which makes the case for openness and open source as the way to AI safety. Kudos on mozilla, we stand by the letter 🤝
Big CO LLMs + APIs
OpenAI - ChatGPT "all tools" aka MMIO mode (that's now dubbed "confidential")
Just a week before the first Dev Day from OpenAI, we were hanging out in X spaces talking about what the regulation might bring, when a few folks noticed that their ChatGPT interface looks different, and saw a very specific popup message saying that ChatGPT can now talk with documents, and "use tools without switching", see and interact with DALL-E and Advanced Data Analysis (FKA Code Interpreter) all in one prompt.
While many X takes focused solely on just how many "chat with your PDF" startups OpenAI just "killed", and indeed, the "work with PDFs" functionality seemed new, chatGPT could now get uploads of files, had the ability to search, to go to a specific page, even do a basic summary on PDF files, I was interested in the second part!
Specifically because given GPT-4V is now basically enabled for everyone, this "combined" mode makes ChatGPT the first MMIO model that we have, which is a multi modal on input (Text, Voice, Images) and output (Text, Images). You see, most MultiModal Models so far have been only multimodal on the input, ie, take in text or images or a combination, and while playing around with the above, we noticed some incredible use-cases that are now available!
ChatGPT (for some lucky folks) can now do all these things in one prompt with shared context:
* Read and interact with PDFs
* See and understand images + text
* Browse & Search up to date info with Bing
* Write and execute code with ADA
* Generate images with DALL-E
All in the same prompt, one after another, and often for several steps and iterations.
One such example was, I asked for "get the current weather in Denver and generate an image based on the conditions" and we got this incredible, almost on the fly "weather" UI, showing the conditions (it was the first snow in CO this year), weather, humidity and everything. Now, DALL-E is ok with text but not great, but it's incredible with scenery, so having this "on the fly UI" that has real time info was super great to show off the capabilities of a general model.
We also saw prompts from folks who uploaded a picture of an obscure object, and asked DALL-E to "add" this object to an already generated image, so DALL-E now has eyes, and can understand and "draw" some of the objects and add them to other images, which was an amazing thing to see, and I can't wait to play around with this functionality.
We noticed a few more things, specifically that DALL-E images are now stored on the same disk that you get access to with ADA, so you can then ask ChatGPT to upscale, crop and do things with those images for example, and generate code with those images as a background!
There are so many new potential use-cases that have opened up, that we spent a long evening / night on X spaces trying to throw the kitchen sink onto this mode, in the fear that it was a fluke by OpenAI and they weren't meant to release this, and we were right! Today on ThursdAI live recording, some users reported that they no longer have access to it (and they miss it!) and some reported that it's now called something like "Confidential"
Someone also leaked the full prompt for this "all tools" mode and it's a doosy! The "All Tools" omni-prompt takes a whopping 2,756 tokens, but it's also using the GPT-4 32k model, with a 32,767 token context window. (link)
I guess we're going to see the announcement on Dev Day (did you set a reminder?)
This new mode that we saw and played with, added to the many many leaks and semi-confirmed modes that are coming out of Reddit make it seem like ChatGPT is going to have an all out Birthday party next week and is about to blow some people's minds! We're here for it! 👏
Open Source LLMs
CodeFusion - 75M parameters model based on Diffusion Model for Code Generation
Code-fusion was taken down from ArXiv, claimed 3.5 is 20B params (and then taken down saying that this was unsubstantiated) X link
The paper itself discusses the ability to use diffusion to generate code, and has much less data to get to a very good coding level, with a model small enough to fit on a chip's cache (not even memory) and be very very fast. Of course, this is only theoretical and we're going to wait for a while until we see if this replicates, especially since the PDF was taken down due to someone attributing the 20B parameters note to a forbes article.
The size of the model, and the performance score on some coding tasks make me very very excited about tiny models on edge/local future!
I find the parameter obsession folks have about OpenAI models incredible, because parameter size really doesn’t matter, it's a bad estimation anyway, OAI can train their models for years and keep them in the same parameter size and they would be vastly different models at the start and finish!
Together releases a massive 30T tokens dataset - RedPajama-Data-v2 (Announcement, HF)
This massive massive dataset is 25x the previous RedPajama, and is completely open, deduplicated and has enormous wealth of data to train models from. For folks who were talking the "there's no more tokens" book, this came as a surprise for sure! It's also multi-lingual, with tokens in English, French, Italian, German and Spanish in there. Kudos to Together compute for this massive massive open source effort 👏
Open source Finetunes Roundup
This week was another crazy one for open source fine-tuners, releasing SOTA after SOTA, many of them on ThursdAI itself 😅 Barely possible to keep up (and that's quite literally my job!)
Mistral 7B - 128K (and 64K) (HF)
The same folks who brought you the YaRN paper, Emozilla, Bowen Peng and Enrico Shippole (frequent friends of the pod, we had quite a few conversations with them in the past) have released the longest context Mistral fine-tune, able to take 64K and a whopping 128K tokens in it's context length, making one of the best open source model now compatible with book length prompts and very very long memory!
Capybara + Obsidian (HF, Quantized)
Friend of the pod (and weekly cohost) LDJ releases 2 Nous research models, Capybara (trained on StableLM 3B and Mistral 7B) and Obsidian, the first vision enabled multi modal 3B model that can run on an iPhone!
Capybara is a great dataset that he compiled and the Obsidian model uses the LLaVa architecture for input multimodality and even shows some understanding of humor in images!
Alignment Lab - OpenChat-3.5 a chatGPT beating open source model (Announcement, HF)
According to friends of the pod Alignment Lab (of OpenOrca fame) we get a Mistral finetune that beats! chatGPT on many code based evaluations (from march, we all think chatGPT became much better since then)
OpenChat is by nature a conversationally focused model optimized to provide a very high quality user experience in addition to performing extremely powerfully on reasoning benchmarks.
Open source is truly unmatched, and in the face of a government regulation week, open sources is coming out in full!
Voice
HuggingFace Distill Whisper - 6x performance of whisper with 1% WER rate (Announcement, HF)
Hugging face folks release a distillation of Whisper, a process (and a paper) with which they use a "teacher" model like the original Open AI whisper, to "teach" a smaller model, and in the process of distillation, transfer capabilities from one to another, while also making the models smaller!
This makes a significantly smaller model (2x smaller) with comparative (and even better) performance on some use-cases, while being 6x faster!
This distill-whisper is now included with latest transformers (and transformers.js) releases and you can start using this faster whisper today! 👏
That's it for today folks, it's been a busy busy week, and many more things were announced, make sure to join our space and if you have read all the way until here, DM me the 🧯 emoji as a reply or in a DM, it’s how I know who are the most engaged users are!
ThursdAI - Recaps of the most high signal AI weekly spaces is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
ThursdAI October 26th
Timestamps and full transcript for your convinience
## [00:00:00] Intro and brief updates
## [00:02:00] Interview with Bo Weng, author of Jina Embeddings V2
## [00:33:40] Hugging Face open sourcing a fast Text Embeddings
## [00:36:52] Data Provenance Initiative at dataprovenance.org
## [00:39:27] LocalLLama effort to compare 39 open source LLMs +
## [00:53:13] Gradio Interview with Abubakar, Xenova, Yuichiro
## [00:56:13] Gradio effects on the open source LLM ecosystem
## [01:02:23] Gradio local URL via Gradio Proxy
## [01:07:10] Local inference on device with Gradio - Lite
## [01:14:02] Transformers.js integration with Gradio-lite
## [01:28:00] Recap and bye bye
Hey everyone, welcome to ThursdAI, this is Alex Volkov, I'm very happy to bring you another weekly installment of 📅 ThursdAI.
ThursdAI - Recaps of the most high signal AI weekly spaces is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
TL;DR of all topics covered:
* Open Source LLMs
* JINA - jina-embeddings-v2 - First OSS embeddings models with 8K context (Announcement, HuggingFace)
* Simon Willison guide to Embeddings (Blogpost)
* Hugging Face - Text embeddings inference (X, Github)
* Data Provenance Initiative - public audit of 1800+ datasets (Announcement)
* Huge open source LLM comparison from r/LocalLLama (Thread)
* Big CO LLMs + APIs
* NVIDIA research new spin on Robot Learning (Announcement, Project)
* Microsoft / Github - Copilot crossed 100 million paying users (X)
* RememberAll open source (X)
* Voice
* Gladia announces multilingual near real time whisper transcriptions (X, Announcement)
* AI Art & Diffusion
* Segmind releases SSD-1B - 50% smaller and 60% faster version of SDXL (Blog, Hugging Face, Demo)
* Prompt techniques
* How to use seeds in DALL-E to add/remove objects from generations (by - Thread)
This week was a mild one in terms of updates, believe it or not, we didn't get a new State of the art open source large language model this week, however, we did get a new state of the art Embeddings model from JinaAI (supporting 8K sequence length).
We also had quite the quiet week from the big dogs, OpenAI is probably sitting on updates until Dev Day (which I'm going to cover for all of you, thanks to Logan for the invite), Google had some leaks about Gemini (we're waiting!) and another AI app builder thing, Apple is teasing new hardware (but nothing AI related) coming soon, and Microsoft / Github announced that CoPilot has 100 million paying users! (I tweeted this and Idan Gazit, Sr. Director GithubNext where Copilot was born, tweeted that "we're literally just getting started" and mentioned November 8th as... a date to watch, so mark your calendars for some craziness next two weeks)
Additionally, we covered the Data provenance initiative that helps sort and validate licenses for over 1800 public datasets, a massive effort led by Shayne Redford with assistance from many folks including friend of the pod Enrico Shippole, we also covered another massive evaluation effort by a user named WolframRavenwolf on the LocalLLama subreddit, that evaluated and compared 39 open source models and GPT4. Not surprisingly the best model right now is the one we covered last week, OpenHermes 7B from Teknium.
Two additional updates were covered, one of them is Gladia AI, released their version of whisper over web-sockets, and I covered it on X with a reaction video, it allows developers to stream speech to text, with very low latency and it's multi-lingual as well, so if you're building an agent that folks can talk to, definitely give this a try, and finally, we covered SegMind SSD-1B, a distilled version of SDXL, making it 50% smaller in size and 60% faster in generation speed (you can play with it here)
This week I was lucky to host 2 deep dive conversations, one with Bo Wang, from Jina AI, and we covered embeddings, vector latent spaces, dimensionality, and how they retrained BERT to allow for longer sequence length, it was a fascinating conversation, even if you don't understand what embeddings are, it's well worth a listen.
And in the second part, I had the pleasure to have Abubakar Abid, head of Gradio at Hugging Face, to talk about Gradio, it's effect on the open source community, and then joined by Yuichiro and Xenova to talk about the next iteration of Gradio, called Gradio-lite that runs completely within the browser, no server required.
A fascinating conversation, if you're a machine learning engineer, AI engineer, or just someone who is interested in this field, we covered a LOT of ground, including Emscripten, python in the browser, Gradio as a tool for ML, webGPU and much more.
I hope you enjoy this deep dive episode with 2 authors of the updates this week, and hope to see you in the next one.
P.S - if you've been participating in the emoji of the week, and have read all the way up to here, your emoji of the week is 🦾, please reply or DM me with it 👀
Timestamps and full transcript for your convinience
## [00:00:00] Intro and brief updates
## [00:02:00] Interview with Bo Weng, author of Jina Embeddings V2
## [00:33:40] Hugging Face open sourcing a fast Text Embeddings
## [00:36:52] Data Provenance Initiative at dataprovenance.org
## [00:39:27] LocalLLama effort to compare 39 open source LLMs +
## [00:53:13] Gradio Interview with Abubakar, Xenova, Yuichiro
## [00:56:13] Gradio effects on the open source LLM ecosystem
## [01:02:23] Gradio local URL via Gradio Proxy
## [01:07:10] Local inference on device with Gradio - Lite
## [01:14:02] Transformers.js integration with Gradio-lite
## [01:28:00] Recap and bye bye
Full Transcription:
[00:00:00] Alex Volkov: Hey, everyone. Welcome to Thursday. My name is Alex Volkov, and I'm very happy to bring you another weekly installment of Thursday. I. This week was actually a mild one in terms of updates, believe it or not. Or we didn't get the new state of the art opensource, large language model this week. However, we did get a new state of the art embeddings model. And we're going to talk about that. we got very lucky that one of the authors of this, a medics model, gold Gina embeddings V2, Bo Wang joined us on stage and gave us a masterclass in embeddings and share some very interesting things about this, including some stuff they haven't charged yet. So definitely worth a listen. Additionally recovered the data provenance initiative that helps sort and validate licenses for over 1800 public data sets. A massive effort led by Shane Redford with assistance from many folks, including a friend of the pod. Enrico Shippole.
[00:01:07] we also covered the massive effort by another user named Wolf from Ravenwolfe on the local Lama subreddit. Uh, that effort evaluated and compared to 39 open source models ranging from 7 billion parameters to 70 billion parameters and threw in the GPT4 comparison as well. Not surprisingly, the best model right now is the one we covered last week from friends of the politic new called open Hermes seven B.
[00:01:34] Do additional updates we've covered. One of them is Gladia AI, a company that offers transcription and translation APIs release their version of whisper over WebSockets. So live transcription, and I covered it on X with a reaction video. And I'll add that link in the show notes. It allows developers like you to stream speech, to text and. Very low latency and high quality and it's multi-lingual as well. So if you're building an agent that your users can talk to. Um, definitely give this a try. And finally Segmind segued mind accompany that just decided to open source a distilled version of. SDXL, making it 50% smaller in size and the in addition to that 60% faster in generation speed. The links to all these will be in the show notes.
[00:02:23] But this week I was lucky to host two deep dives, one with Bo Weng which I mentioned. Uh, we've covered the embeddings vector led in spaces that dimensionality and how they retrained Bert model to allow for a longer sequence length. It was a fascinating conversation. Even if you don't understand what embeddings are, it's well worth the listen. And, , I learned a lot. Now I hope you will, as well. And the second part, I had the pleasure to have a Brubaker a bit. The head of grandio at hugging face to talk about gradient. What is it? Uh, its effect on the open source community. And then joined by utero. And Sunnova to talk about the next iteration of Grigio called Grigio light that runs completely within the browser. No Serra required. We also covered a bit of what's coming to Gradio in the next release. on October 31st.
[00:03:15] A fascinating conversation. If you're a machine learning engineer, AI engineer, or just somebody who's interested in this skilled. You've probably used radio, even if you haven't written any Gradio apps, every model and hugging face usually gets a great deal demo.
[00:03:30] And we've covered a lot of ground, including M scripting. Then by filling the browser. As a tool for machine learning, web GPU, and so much more.
[00:03:38] Again, fascinating conversation. I hope you enjoy this deep dive episode. Um, humbled by the fact that sometimes the people. Who produced the updates we cover actually come to Thursday and talk to me about the things they released. And I hope this trend continues, and I hope you enjoyed this deep dive over an episode. And, um, I'll see you in the next one. And now I give you thursday october 26. oh, awesome. It looks like Bo, you joined us. Let's see if you're connecting to the audience, and can you unmute yourself, can you see if we can hear you?
[00:04:22] Bo Wang: Hi, can you hear me? Oh, we can hear you fine, awesome. this, this, this feature of, of Twitter.
[00:04:30] Alex Volkov: That's awesome. This, this usually happens, folks join and it's their first face and then they can't leave us. And so let me just do a little, maybe... Maybe, actually, maybe you can do it, right? Let me just present yourself.
[00:04:42] I think I followed you a while ago, because I've been mentioning embeddings and the MTB dashboard and Hug and Face for a while. And, obviously, embeddings are not a new concept, right? We started with Word2Vec ten years ago, but now, with the rise of LLMs, And now with the rise of AI tools and many people wanting to understand the similarity between the user query and an actual thing they, they, they stored in some database, embeddings have seen a huge boon.
[00:05:10] And also we've saw like all the vector databases pop up like mushrooms after the rain. I think Spotify just released a new one. And my tweet was like, Hey, do we really need another vector database? But Boaz, I think I started following you because you mentioned that you were working on something that's.
[00:05:25] It's coming very soon, and finally this week this was released. So actually, thank you for joining us, Beau, and thank you for doing the first ever Twitter space for yourself. How about can we start with your introduction of who you are and how are you involved with this effort, and then we can talk about Jina.
[00:05:41] Bo Wang: Yes, sure. Basically I have a very different background. I guess I was oriJinally from China, but my bachelor was more related to text retrieval. I have a retrieval experience rather than pure machine learning background, I would say. Then I came to the Europe. I came to the Netherlands like seven or eight years ago as a, as an international student.
[00:06:04] And I was really, really lucky and met my supervisor there. She basically guided me into the, in the world of the multimedia information retrieval, multimodal information retrieval, this kind of thing. And that was around 2015 or 2016. So I also picked up machine learning there because when I was doing my bachelor, it's not really hot at that moment.
[00:06:27] It's like 2013, 2014. Then machine learning becomes really good. And then I was really motivated, okay, how can I apply machine learning to, to search? That is, that is my biggest motivation. So when I was doing my master, I, I collaborated with my friends in, in, in the US, in China, in Europe. We started with a project called Match Zoo.
[00:06:51] And at that time, the embedding on search is just a nothing. We basically built a open source. Software and became at that time the standard of neural retrieval or neural search, this kind of thing. Then when the bird got released, then our project basically got queue because. Everyone's focus basically shifted to BERT, but it's quite interesting.
[00:07:16] Then I graduated and started to work as a machine learning engineer for three years in Amsterdam. Then I moved to Berlin and joined Jina AI three years ago as a machine learning engineer. Then basically always doing neural search, vector search, how to use machine learning to improve search. That is my biggest motivation.
[00:07:37] That's it.
[00:07:38] Alex Volkov: Awesome. Thank you. And thank you for sharing with us and, and coming up and Gene. ai is the company that you're now working and the embeddings thing that we're going to talk about is from Gene. ai. I will just mention the one thing that I missed in my introduction is the reason why embeddings are so hot right now.
[00:07:53] The reason why vectorDB is so hot right now is that pretty much everybody does RAG, Retrieval Augmented Generation. And obviously, For that, you have to store some information in embeddings, you have to do some retrieval, you have to figure out how to do chunking of your text, you have to figure out how to do the retrieval, like all these things.
[00:08:10] Many people understand that whether or not in context learning is this incredible thing for LLMs, and you can do a lot with it, you may not want to spend as much tokens on your allowance, right? Or you maybe not have enough in the context window in some in some other LLMs. So embeddings... Are a way for us to do one of the main ways to interact with these models right now, which is RAC.
[00:08:33] And I think we've covered open source embeddings compared to OpenAI's ADA002 embedding model a while ago, on ThursDAI. And I think It's been clear that models like GTE and BGE, I think those are the top ones, at least before you guys released, on the Hugging Face big embedding model kind of leaderboard, and thank you Hugging Face for doing this leaderboard.
[00:09:02] They are great for open source, but I think recently it was talked about they're lacking some context. And Bo, if you don't mind, please present what you guys open sourced this week, or released this week, I guess it's open source as well. Please talk through Jina Embeddings v2 and how it differs from everything else we've talked about.
[00:09:21] Bo Wang: Okay, good. Basically, it's not like embeddings for, how can I say, maybe two... point five years. But previously we are doing at a much smaller scale. Basically we built all the algorithm, all the platform, even like cloud fine tuning platform to helping people build better embeddings. So there is a not really open source, but a closed source project called fine tuner, which we built to helping user build better embeddings.
[00:09:53] But we didn't, we found it okay. Maybe we are maybe too early. because people are not even using embeddings. How could they find embeddings? So we decided to make a move. Basically, we basically scaled up our how can I say ambition. We decided to train, train our own embeddings. So six months ago, we started to train from scratch, but not really from scratch because in binding training, normally you have to train in two stages.
[00:10:23] The first stage, you need to pre train on massive scale of like text pairs. Your objective is to bring these text pairs as closer as possible, as possible, because these text pairs should be semantically related to each other. In the next stage, you need to fine tune with Carefully selected triplets, all this kind of thing.
[00:10:43] So we basically started from scratch, but by collecting data, I think it was like six months ago, we working with three to four engineers together, basically scouting every possible pairs from the internet. Then we basically created like one billion, 1. 2 billion sentence pairs from there. And we started to train our model based on the T5.
[00:11:07] Basically it's a very popular encoder decoder model. This is on the market. But if you look at the MTB leaderboard or all the models on the market, the reason why they only support 512 sequence lengths is constrained actually by the backbone itself. Okay, we figure out another reason after we release the V1 model.
[00:11:31] Basically, if you look at. And the leaderboard or massive text embedding leaderboard, that is the one Alex just mentioned. Sorry, it's really bad because everyone is trying to overfitting the leaderboard. That naturally happens because if you look at BGE, GTE, the scores will never that high if you don't add the training data into the, into the, That's really bad.
[00:12:00] And we decided to take a different approach. Okay. The biggest problem we want to solve first, improving the quality of the embeddings. The second thing we want to solve is. Enable user to making longer context lens. If we want to making user make user have longer context lens, so we have to rework the BERT model, because every basically the embedding model, the backbone was from BERT or T5.
[00:12:27] So we basically started from scratch. Why not we just borrow the latest research from large language model? Every large language model wants large context. Why not we just borrow the research ideas? into the musk language modeling modelings. So we basically borrowed some ideas, such as rotary position embeddings or alibi, maybe you did, and reworked BERT.
[00:12:49] We call it JinaBERT. So basically now the JinaBERT can handle much longer sequence. So we trained BERT from scratch. Now BERT has been a byproduct of our embeddings. Then we use this JinaBERT to contrastively train the models on the semantic pairs and triplets that finally allow us to encode 8K content.
[00:13:15] Alex Volkov: Wow, that's impressive. Just, just to react to what you're saying, because BERT is pretty much every, everyone uses BERT or at least use BERT, right? At least in the MTB leaderboard. I've also noticed many other examples that use BERT or distilled BERT and stuff like this. You're saying, what you're saying, if I'm understanding correctly, is this was the limitation for sequence length?
[00:13:36] for other embedding models in the open source, right? And the OpenAI one that's not open source, that does have 8, 000 sequence length. Basically, sequence length, if I'm explaining correctly, is just how much text you can embed without chunking.
[00:13:51] Yes. And you're basically saying that you, you guys saw this limitation and then retrained BERT to use rotary embeddings. We've talked about rotary embeddings multiple times here. We had folks behind the yarn paper for extending context windows. Alibi is we follow Ophir Press.
[00:14:08] I don't think Ophir ever joined ThursdAI, but Ophir, if you hear this, you're welcome to join as well. So Alibi is another way to extend context windows and I think Mosaic folks used Alibha and some other folks as well. Bo, could you speak more about like borrowing the context from there and retraining BERT to JinaBERT and whether or not JinaBERT is also open source?
[00:14:28] Bo Wang: Oh, we actually want to make JinaBERT open source, but I need to align with my colleagues. That's, that's, that's really, that's a decision to be made. And the, the idea is quite naive. If you didn't know, I don't want to dive into too much about technical details, but basically the idea of Alibi basically removed the position embeddings from the large language model pre training.
[00:14:55] And the Alibi technique allow us to train on the shorter sequences. But inference at every very long sequence. So in the end, I think if I, my remember is correct, the author of alibi paper, basically trained model on 512 sequence lens and 1,024 sequence lens, but he's able to inference on 16 K. 16 K, like sequence lens.
[00:15:23] If you further expand it, you are not capable because that's the limitation of hardware, that's the limitation of GPE. So he, he actually tested 16 K like a sequence lens. So what we did is just. Borrowed this idea from the autoregressive models into the mask language models. And integrate Alibi, remove the position embeddings from the bird, and add this Alibi slope and all the Alibi stuff back into the bird.
[00:15:49] And just borrowed the things how we train bird or something Roberta, something from Roberta, and retrained the bird. I never imagined bird could be a by product of our embedding model, but this... This happened. We could open source it. Maybe I have to discuss with my colleague.
[00:16:09] Alex Volkov: Okay. So when you talk to your colleagues, tell them that first of all, you already said that you may do this on ThursdAI Stage.
[00:16:15] So your colleagues are welcome also to join. And when you open source this, you guys are welcome to come here and tell us about this. We love the open source. The more you guys do, the better. And the more it happens on ThursdAI Stage, the better, of course, as well. Bo, you guys released the Jina Embedding Version 2, correct?
[00:16:33] Gene Embedding Version 2 has a sequence length of 8k tokens. So that actually allows to, if, just for folks in the audience, 8, 000 tokens is, I want to say, maybe like 6, 000 words in English around, right? And different languages as well. Could you talk about multilinguality as well? Is it multilingual, is it only English?
[00:16:53] How that how that appears within the embedding model?
[00:16:57] Bo Wang: Okay, actually, our Jina Embedding V2 is only English, so it's a monolingual embedding model. If you look at the MTV benchmark or all the public multilingual models, they are multilingual. But to be frankly, I don't think this is a fair solution for that.
[00:17:18] I think at least every major language.
[00:17:24] We decided to choose another hard way. We will not train a multilingual model, but we will train a bilingual model. Our first target will be German and Spanish. What we are doing at Jina AI is we basically Fix our English embedding model as it is just keep it at is, but we are continuously adding the German data, adding the Spanish data into the embedding model.
[00:17:51] And our embedding model cares two things. We make it bilingual. So it's either German, English or German English, Spanish, Spanish, English, German, English, or Japanese, English, whatever. And what we are doing is we want to build this embedding model to make it monolingual. So imagine you are, you have a German English embedding model.
[00:18:12] So if you search for German, you'll get German results. If you use English, you'll get English results. But we also care about the cross linguality of this bilingual model. So imagine you, you, you encode two, two sentences. One is in German, one is in English, which they are With the same meaning, we also want these vectors to be mapped into the similar semantic space.
[00:18:36] Because I, I'm a foreigner myself, sometimes, imagine I, I, I buy some stuff in the supermarket. Sometimes I have to translate, use Google Translate, for example, milk into Milch in German, then, then, then put it into the search box. I really want this bilingual model happen. And I believe every, at least, major language deserves such an embedding model.
[00:19:03] Alex Volkov: Absolutely. And thanks for clarifying this because one of the things that I often talk about here on Thursday Night is as a founder of Targum, which is inside videos, is just how much language barriers are preventing folks from conversing to each other. And definitely embeddings are... The way people extend memories parallel lines, right?
[00:19:21] So like a huge, a huge thing that you guys are working on and especially helpful. The sequence length is, and I think we have a question from the audience is what is the sequence lengths actually allow people to do? I guess Jina and I worked with some, some other folks in the embedding space. Could you talk about what is the longer sequence lengths now unlocking for people who want to use open source embeddings?
[00:19:41] Obviously. My answer here is, well, OpenAI's embeddings is the one that's most widely used, but that one you have to do online, and you have to send it to OpenAI, you have to have a credit card with them, blah, blah, blah, you have to be from supported countries. Could you talk about a little bit of what sequence length allows unlocks once you guys release something like this?
[00:20:02] Bo Wang: Okay, actually, we didn't think too much about applications. Most of the vector embeddings applications, you can imagine search and classification. You build another layer of, I don't know, classifier to classify items based on the representation. You can build some clustering. You can do some anomaly detection on the NLP text.
[00:20:22] This is something I can imagine. But the most important thing I I have to be frankly to you because we are, we are like writing a technical report as well. Something like a paper maybe we'll submit to academic conference. Longer embeddings doesn't really always work. That is because sometimes if the important message is in in the front of the document you want to embed, then it makes most of the sense just to encode let's say 256 tokens.
[00:20:53] or 512. But sometimes if you you have a document which the answer is at the middle or the end of the document, then you will never find it if the message is truncated. Another situation we find very interesting is for clustering tasks. Imagine you want to visualize your embeddings. Longer longer sequence length almost always helps and for clustering tasks.
[00:21:21] And to be frankly, I don't care too much about the application. I think people, we, what we're offering is the, how can I say, offering is, is like a key. We, we unlock this 512 sequence length. To educate and people can explore it. People, let's say I, I only need two K then, then people just set tokenize max lens to two k.
[00:21:44] Then, then embed. Based on their needs, I just don't want to be, people to be limited by the backbone, by the 500 to 12 sequence lengths. I think that's the most important thing.
[00:21:55] Alex Volkov: That's awesome. Thank you. Thank you for that. Thank you for your honesty as well. I love it. I appreciate it. The fact that, there's research and there's application and you not necessarily have to be limited with the application set in mind.
[00:22:07] We do research because you're just opening up doors. And I love, I love hearing that. Bo maybe last thing that I would love to talk to you about as the expert here on the topic of dimensions. Right. So dimensionality with embeddings I think is very important. Open the eye, I think is one of the highest ones.
[00:22:21] The kind of the, the thing that they give us is like 1200 mentioned as well. You guys, I
[00:22:26] think
[00:22:26] Jina is around 500. Or so is that correct? Could you talk a bit about that concept in broad strokes for people who may be not familiar? And then also talk about the why the state of the art OpenAI is so far ahead?
[00:22:39] And what will it take to get the open source embeddings also to catch up in dimensionality?
[00:22:46] Bo Wang: You mean the dimensionality of the vectors? Okay, basically we follow a very standard BERT size. The only thing we modified is actually the the alibi part and some training part.
[00:22:58] And our small model dimensionality is 512, and the base model is 768 and we have also a large model, haven't been released because of the training is too slow. We have so much data to change. Even the model size is small, but we have so much data and so large model dimensionality size is 1,024. And if my memory is correct, so are I embedding 0 0 2?
[00:23:23] Have but dimensionality of. 1, 5, 3, 6, something like that, which is a very strange dimensionality, I have to say, but I would say the dimensionality is, is, is the longer might be more Better or more expressive, but shorter, which means when you are doing the vector search, it's gonna be much more faster.
[00:23:48] So it's something you have to balance. So if you think the speed query speed, or the retrieval speed or whatever is more important to you. And if I, if I know correct, some of the Vector database, they make money by the dimensionality, let's say. They, they charge you by the dimensionality, so it's actually quite expensive if your dimensionality is too high.
[00:24:13] So it's a balance between expressionist and the, the, the, the speed and the, the, the, the cost you want to invest. So it's. It's very hard to determine, but I think 512, 768, and 1024 is very common as BERT.
[00:24:34] Alex Volkov: So great to hear that a bigger model is also coming, but it hasn't been released yet. So there's like the base model and the small model for embeddings, and we're waiting for the next one as well.
[00:24:46] I wanted to maybe ask you to maybe simplify for the audience, the concept of dimensionality. What does it mean between, what is the difference between embeddings that were stored with 512 and like 1235 or whatever OpenAI does? What does it mean for quality? So you mentioned the speed, right? It's easier to look up nearest neighbors, maybe within the 512 dimension space, what does it actually mean for quality of look up of different other ways that strings can compare? Could you maybe simplify the whole concept, if possible, for people who don't speak embeddings?
[00:25:19] Bo Wang: Okay maybe let me quickly start with the most basic version.
[00:25:24] If you imagine, if you type something in the search box right now, when doing, doing the matching, and it's actually also embedding, but it's something like if I make a simple version, it's a binary embedding. Imagine there 3, 000 words in English. Maybe there are much more, definitely. Imagine it's 3, 000 words in English, then the vector is 3, 000 dimensionality.
[00:25:48] Then what current solution of searching or matching do is just making... If the query has a token, if your document has a token, if your document has this token, then your occurrence will be one. If you query has the token, and this one will match your document token. But it's also about the, the frequency it appears, it's how, how rare it is.
[00:26:12] But the current solution is basically matching by the. By the English word, but with neural network, basically if you know about this, for example, ResNet know about a lot of different, for example, classification models, basically the output class of item, but if you chop up the classification layer, it will give you some a vector.
[00:26:36] Basically this vector is It's the representation of the information you want to encode. Basically it's a compressed version of the information in a certain dimensionality such as 512, 768, something like this. So it's a compressed list of non numerical numbers, which we normally call it dense vectors.
[00:26:57] because it's much more how can I say in English dense, right? Compared to the traditional way we store vectors, it's much more sparse. There is a lot of zero, there is a lot of one, because zero means not exist, one means exist. When one exists, then there is a match, then you've got the search result.
[00:27:16] So these dense vectors capture more about semantics, but if you match by the occurrence, then you might lose the semantics. But only matching by the occurrence of a token or a word.
[00:27:31] Alex Volkov: Thank you. More dimensions, basically, if I'm not saying it correctly, more dimensions just have more similarity vector. So like more things two strings or tokens can be similar on. And this basically means higher match rate. For more similarity things. And I think the basic stuff I think is covered in the Simon Wilson, the first pin tweet here, Simon Wilson did a basic, basic intro into what do dimensions embeddings mean and why they matter.
[00:28:00] And I specifically love the fact that there's arithmetic that can be done. I think somebody reads the paper even before this whole LLM thing, where if you take embeddings for Spain and embeddings for Germany, and then you take you, you can subtract like the embedding for Paris and then you get something closer to, to like Berlin, for example, right?
[00:28:19] So there's like concepts in, inside these things that are they're even arithmetic works and if you take like King and you subscribe male, then you get something closer to Queen and stuff like this. It's really, really interesting. And also Bo, you mentioned visualization as well. It's really impossible to visualize.
[00:28:36] 10, 24, et cetera, dimensions, right? Like we humans, we have perceived maybe three, maybe three and a half, four with time, whatever. And usually what happens is those multiple dimensions get down scaled to 3D in order to visualize in neighborhoods. And I think we've talked with folks from ARISE. They have a software called Phoenix that allows you to visualize embeddings for clustering and for semantics.
[00:29:02] Atlas does this as well, right? Nomic AI's Atlas does this as well. You can provide dimensions as well. And so you can provide embeddings and see clustering for concepts. And it's really pretty cool. If you haven't played with this, if you only did VectorDBs and you stored your stuff after you've done chunking, but you've never visualized how this looks, I strongly recommend you to do and I think well, thank you so much for joining us and explaining to us, the internals and sharing with us some exciting things about what's to come. Jina Burt is hopefully hopefully is coming, a, a retrained version of Burt, the, the, the, the, the... The grease of all how should I say, I can't, it's hard for me to define a verb, but I see it everywhere it's, it's the big base bone of a lot of NLP tasks, and it's great to see that you guys are about to first of all, retrain it for longer sequences, using tricks like Alibi and and I think you said Positional Embeddings, and hoping to see some open source action from this, but also that Jina Embedding's large model is coming as well with more dimensions waiting for that. Hopefully you guys didn't stop training that. And I just want to tell folks why I'm excited for this. And this kind of will take us to the next.
[00:30:08] Point as well is because, while I love OpenAI, I honestly do, I'm going to join their Dev Day, I'm going to report from their Dev Day and tell you all the interesting things that OpenAI does. We've been talking about we've been talking and we'll be talking today about local inference, about running models on edge, about running models of your own.
[00:30:28] Mistin is here, he even works on some bootable stuff that you can like completely off the grid run. And, so far, we've been focused on open source LLMs, for example, right? So we've had I see Pharrell in the audience from Skunks Works, and many other fine tuners, like Tignium, Alignment Labs, all these folks are working on local LLMs, and they never get to GPT 4 level yet.
[00:30:51] We're waiting for that, and they will. But the whole point of them is, you run them locally, they're uncensored, you can do whatever you want, you can fine tune them on whatever you want. However, the kind of the embeddings part Is the glue to connect it to an application and the reason is because there's only so much context window also context window is expensive and even if theoretically the yarn paper that we've talked with the authors of allows you to extend the context window to 128, 000 tokens The hardware requirements for that are incredible, right?
[00:31:22] Everybody in the world of AI engineers, they switch up to, to, to retrieval of data generation. Basically, instead of shoving everything in the context, they switched Hey, let's use a vector database. Let's say a Chroma. Or Pinecone, or Waviate, like all of those, vectorized from Cloudflare, and the other one from Spotify there, I forget its name or even Superbase now has one.
[00:31:43] Everybody has a vector database it seems these days, and the reason for that is because all the AI engineers now understand that you need to put some text into some embeddings, store them in some database. And many pieces of that were still requiring internet, requiring OpenAI API calls, requiring credit cards, like all these things.
[00:32:03] And I think it's great that we've finally got to a point where, first of all there are embeddings that are matching whatever OpenAI has given us. And now you can run them locally as well. You don't have to go to OpenAI. If you don't want to host, you can probably run them. I think though GeneEmbedding's base is very tiny.
[00:32:20] Like it's half like the small model is 770 megabytes, I think. Maybe a little bit more, if
[00:32:27] Bo Wang: I'm looking at this correctly. Sorry, it's half precision. So you need to double it to make it FP32.
[00:32:33] Alex Volkov: Oh yeah, it's half precision. So it's already quantized, you mean?
[00:32:37] Bo Wang: Oh no, it's just to store it as FV16,
[00:32:39] Alex Volkov: if you store it as FV16.
[00:32:43] Oh, if you store it as FV16. But the whole point is the next segment in ThursdAI today is going to be less about updates and more about the very specific things. We've been talking about local inference as well, and these models are tiny, you can run them on your own hardware, on Edge via Cloudflare, let's say, or on your computer.
[00:32:58] And you now can do almost end to end application wise. From the point of your user inputting a query embedding this query, running a match, a vector search, KNNN and whatever you want nearest neighbor search for that query for the user. Retrieve that all from like local open source. You basically you, you can basically go offline.
[00:33:20] And this is what we want in, in the era of upcoming regulation towards what AI can be and cannot be. And the era of like open source models getting better and better. We've talked last week where Zephyr and I think Mistral News from Technium is also matching some GPT 3. 5. All of those models you can download and nobody can tell you not to run inference on them.
[00:33:40] Hugging Face open sourcing a fast Text Embeddings Inference Server with Rust / Candle
[00:33:40] Alex Volkov: But the actual applications, they still require the web or they used to. And now I'm, I'm loving this like new move towards. Even the application layer, even the RAG systems, which are augmented generation, even the vector databases, and even the embeddings are now coming to, to open source, coming to your local computer.
[00:33:57] And this will just mean like more applications either on your phone or your computer. And absolutely love that. Bo, thank you for that. And thank you for coming to the stage here and talking about the things that you guys open sourced and hopefully we'll see more open source from Jina and everybody should follow you and, and Jina as well.
[00:34:13] Thank you. It looks like. Thank you for joining. I think the next thing that I wanna talk about is actually in this vein as well. Let me go find this o Of course, we love hug and face and the thing that I think that's already on top if you look, yeah if you look at the last thing, last tweet that's pinned it's a tweet from Jeri Lou from Lama Index, obviously.
[00:34:33] Well, well, well, we're following Jerry and whatever they're building and doing over at Lama Index because they implement everything like super fast. I think they also added support for Jina like extremely fast. He talks about this thing where HugInFace opensource for us something in Rust and Candlestick?
[00:34:51] Candlelight? Something like that? I forgot that they're like iteration on top of Rust. Basically, the open source is a server that's called TextEmbeddingsInferenceServer that you can run on your hardware, on your Linux. boxes and basically get the same thing that you get from OpenAI Embeddings.
[00:35:07] Because Embeddings is just one thing, but it's a model. And I think you could use this model. You could use this model with transformers but it wasn't as fast. And as Bo previously mentioned, there's considerations of latency for user experience, right? If you're building an application, you want it to be as responsive as possible.
[00:35:24] You need to look at all the places in your stack and say, Hey. What slows me down? For many of us, the actual inference, let's say use GPT 4, waiting on OpenAI to respond and stream that response is what slows many applications down. And but many people who do embeddings, let's say you have a interface of a chat or a search, you need to embed every query the user sends to you.
[00:35:48] And one such slowness there is how do you actually How do you actually embed this? And so it's great to see that Hackenface is working on that and improving that. So you previously could do this with transformers, and now they released this specific server for embeddings called TextEmbeddings Inference Server.
[00:36:04] And I think it's four, four times faster. than the previous way to run this, and I absolutely love it. So I wanted to highlight this in case you are interested. You don't have to, you can use OpenAI Embeddings. Like we said, we love OpenAI, it's very cheap. But if you are interested in doing the local embedding way, if you want to go end to end, complete, like offline, you want to build like an offline application, using their internet server I think is a good idea.
[00:36:29] And also it shows what HuggingFace is doing with Rust and I really need to remember what language there is but definitely a great attempt from Hug and Face, and yeah, just wanted to highlight that. Let's see. Before we are joined from the Grad. io folks, and I think there's some folks in the audience who are ready from Grad.
[00:36:48] io to come up here and talk about local inference which 15 minutes left,
[00:36:52] Data Provenance Initiative at dataprovenance.org
[00:36:52] Alex Volkov: I wanted to also mention the Data Provenance Initiative. Let me actually find this announcement, and then quickly... Quickly paste this here, and I was hoping that Enrico can be here. . There's a guy named Shane Longfree,
[00:37:05] and he released this massive, massive effort, included with many people. And basically what this effort is, it's called the Data Provenance Initiative. Data Provenance Initiative is now existing in dataprovenance. org. And hopefully can somebody maybe send me the, the direct link to the suite to add this.
[00:37:23] It... It is a massive effort to take 1, 800, so 1, 800 Instruct and Align datasets that are public, and to go through them to identify multiple things. You can filter them, exclude them, you can look at creators, and the most important thing, you can look at licenses. Why would you do this? Well, I don't know if somebody who builds an application needs this necessarily, but everybody who wants to fine tune models, the data is the most important key for this, and building data sets and running them through your fine tuning efforts is basically the number one thing that many people do in the fine tune community, right?
[00:38:04] Data wranglers, and now, thank you, Nishtan, thank you so much, and a friend of the pod, Enrico. is now pinned to the top of the tweet. Thank you for to the top of the space, the nest, whatever it's called. A friend of Enrico Cipolla, who we've talked previously in the context of extending I think Lama to first 16k and then 128k.
[00:38:24] I think Enrico is part of the team on yarn paper as well. I joined this effort, and I was hoping Enrique could join us to talk about this. But basically, if you're doing anything with data, this seems like a massive, massive effort. Many datasets from Lion, and we've talked about Lion, and Alpaca, GPT 4L Gorilla, all these datasets.
[00:38:46] It's very important when you release your model as open source that you have the license to actually release this. You don't want to get exposure, you don't want to get sued, whatever. And if you're in finding data sets and creating different mixes to fine tune different models, this is a very important thing.
[00:39:03] And we want to shout out, Shane Longpre, Enrico, and everybody who worked on this because I think... Just, I love these efforts for the open source, for the community, and it just makes, it's easier to fine tune, to train models. It makes it easier for us to advance and get better and smaller models, and it's worth celebrating and ThursdAI is the place to celebrate this, right?
[00:39:27] LocalLLama effort to compare 39 open source LLMs + GPT4
[00:39:27] Alex Volkov: On the topic of extreme, how should I say efforts that are happening by the community on the same topic, I want to add another one, and this one I think I have a way to pull it up, so give me just a second give me just a second, yes. A Twitter user named Wolfram Ravenwolf who is a participant of the local Lama community on Reddit and now is pinned to the nest at the top of the tweet did this massive effort of comparing open source LLMs and tested 39 different models ranging from 7 billion parameters to 70 billion, and also compared them to chat GPT, GPT 4.
[00:40:06] And I just want to circle back to something we've said. In the previous space as well, and I welcome like folks on stage also to jam in here. I've seen also the same kind of concepts from Hug and Face folks. I think Glenn said the same thing. It's really unfair to, to take a open source model like Mistral7b and then start comparing this to GPT 4.
[00:40:26] It's unfair for several reasons. But also I think it, it, it can obfuscate to some people when they do this comparison of how, just how advanced we've come for the past year in open source models. OpenAI has the infrastructure, they're backed by Microsoft, they have the pipelines to serve these models way faster.
[00:40:47] And also, those models don't run on like local hardware, they don't run on like one GPU. It's like a whole, a whole... Amazing MLOps effort to bring you this speed. When you're running local source model open source models locally when they're open source, they're, they're, they're small, there's drawbacks and there's like takeaways that you have to bake in into your evaluation.
[00:41:09] So comparing to the GPT 4, which is super general in many, many things, that will just lead to your disappointment. However, and we've been talking about this like with other open source models If you have a different benchmark in your head of if you're comparing open source to open source, then it's a whole completely different ballgame.
[00:41:26] And then you start seeing things like, Hey, we're noticing that the 7 billion parameter model is, beating 70 billion. We're noticing that size is not necessarily the king because if you guys remember, Three months ago, ni, I wanna say we've talked about Falcon 180 B. 180 B was like, three times the size of like the, the, the next largest model.
[00:41:47] And it was incredible the Falcon Open source this, and then it was like, like a wo like, no, nobody really was able to run 180 B because it's huge. But also once, once we did run it, we saw that like the difference between that and LAMA are not great at all. Maybe a few percentage points on, on the valuations.
[00:42:04] However, the benefits that we see are from local, like for tinier and tinier models from like 7D Mistral, for example, which is the, the one that the fine tuners of the world are now preferring to everything else. And so the kind of, when you're about to evaluate whatever next model that's coming up that we're going to talk about please remember that Comparing to large, open, big companies backed by billions of dollars that run on multiple split hardware, it's just going to lead to disappointment.
[00:42:34] However, when you do some comparisons, like the guy did, that is now pinned to the tweet this is the way to actually do this. However, on specific tasks, like for, say, coding go ahead,
[00:42:46] Nisten Tahiraj: Nisam. I was going to say, we're still a bit early to judge for example, Falcon could have used a lot more training.
[00:42:53] There's also other. parts where larger models play a big effect stuff like if you want to do very long context summarization then you want to use the 70b and as far as i'm getting it and this is probably inaccurate right now but the more tokens you have the more meat you have in there the Then the larger the thoughts can can be.
[00:43:23] So that's the principle which are going by. Well, Mistral will do extremely well in small analytical tasks and in benchmarks, and it's amazing as a tool. It doesn't necessarily mean that it'll be good at thinking big. You still need The meat there, the amount of tokens to do that. Now you could chop it up and, and do it one, one at a time, but anyway, just something to keep in mind because lately we also saw the announcement from Lama70Blong, which started getting really good at at summarization.
[00:44:04] So again, there's one particular part. Which is summarization where you, it looks like you need longer you need bigger models for that. And I've tested it myself with Falcon and stuff, and it's pretty good at summarization, I just want to also give them the benefit of the doubt that there is still something that could be done there.
[00:44:28] I wouldn't just outright dismiss.
[00:44:31] Alex Volkov: Yeah, absolutely, absolutely, and I want to join this non dismissal. Falcon open sourced fully commercially like Falcon 70B before, and this was the biggest open source model at the time. And then they gave us 180B, they didn't have to, and we appreciate like open sourcing.
[00:44:46] We're not going to say no. Bigger models have more information, more, more, maybe world model in them, and there's definitely place for that, for sure. The, the next thing you mentioned also, and I think I, I strongly connect to that, Nissen, and thank you, is GPT 4, for example, is very generalized. It does many, many, many things well.
[00:45:08] It's like kind of impressive and whatever Gemini is going to be from Google soon, hopefully, we're always waiting on ThursdAI, that the breaking news will come on Thursday. We're gonna be talking about something else and then Google suddenly drops Gemini on us. There's also other rumors for Google's other stuff.
[00:45:22] Whatever OpenAI's Arrakis was, and then they stopped training, and whatever next they're coming from OpenAI, will probably blow everything we expect in terms of generality out of the water. And this, the open source models, as, as they currently are, they're really great at... Focused tasks, right? So like the coder model, for example that recently Glaive Coder was released by Anton Bakaj, I think is doing very well on the evaluations for code.
[00:45:51] However, on general stuff, it's probably less, less good. And I think for open models expecting generality on the same level as GPT 4, I think, is, is going to lead to disappointment. But for tasks, I see, I think we're coming close to different things that a year ago seemed state of the art. If you guys remember, it's not even a year since JetGPT was released, right?
[00:46:14] I think JetGPT was released in November? No, not as an API even, just the UI, like middle of November. So we're coming up on one year, I think the Dev Day will actually be one year. That was 3. 5. 3. 5 now, many people use 3. 5 for applications, but, you want to go for 4. If you're paying for Chattopadhyay Plus and you have a task to solve, you're not going to go 3.
[00:46:35] 5 just because you feel like it. You know that 4 is better. But now we're having open source models way smaller. They're actually getting to some levels of 3. 5, and the above effort is actually an effort to try to figure out which ones. And so I strongly recommend, first of all, to get familiar with local Llama subreddit.
[00:46:54] If you don't use Reddit, I feel you, and I've been a Reddit user for a long time, and I stopped. Some parts of Reddit are really annoying. This is actually a very good one, where I get a bunch of my information outside of Twitter. And I think Andrej Karpathy also recommended this recently, which... Then became an item on that subreddit.
[00:47:12] It was really funny. And this massive effort was done by this user and he, he did like a full comparison of just 39 different models. And he outlined the testing methodology as well. We've talked about testing and evaluation methodology. Between ourselves, it's not easy to evaluate these things. A lot of them are like gut feeling.
[00:47:31] A lot of the, the evaluation, and Nathan and I have like our own prompts that we try on every new model, right? It's, it's like a lot, a lot of this for many people is like gut feel. And many people also talk about the problem with evals and I think Bo mentioned the same thing with the embedding leaderboards that then, you know.
[00:47:48] It then becomes like a sport for people to like fine tune and, and release models just to put their name on the board to overfit on whatever whatever metrics and evaluations happen there. And then there's a whole discussion on Twitter whether or not this new model that beats that model on, on some, some score actually, was trained on the actual evaluation data.
[00:48:09] But... Definitely the gut feeling variation is important and definitely having different things to test for is important. And you guys know, I think, those of you who come to ThursdAI, my specific gut feels are about like translation and multilingual abilities, for example, and direction following some other people like Jeremy, Jeremy Howard from ThursdAI have his own like approach.
[00:48:29] Everybody has their own approach. I think what's interesting there is... Kind of the community provides, right? We're like this decentralized brain of evaluating every new model. And for now, the community definitely landed on Mistral as being like the top. At least a model in the 7b range, and Falcon, even though it's huge and can do some tasks like Nissan said is less, less and Lama was there before. So if you start measuring the community responses to open source models, you start noticing better what does what. And this effort from this guy, he actually outlined the methodology, and I want to shout out... Friends of the pod, Tignium being the go to many, many things, specifically because Open Hermes, which, Hermes, which we've talked about before which was fine tuned from Mistral7b is probably like getting the, the, the top leaderboard from there, but also based on my experiences, right?
[00:49:22] So we've talked last week about Open Hermes being able, you're able to run Open Hermes on your... Basically, M1, M2, Max with LM Studio, which also shout out to LM Studio, they're great, and I've tested this, and this seems to be, like, a very, very well rounded model, especially for one that you can run yourself and comparing to GPT 4 and other stuff, this model for specific things is really great.
[00:49:45] It's good for coding. It's not the best for coding. I think there's a coding equivalent. And I just really encourage you, if you're interested, like figuring out what to use. And we've talked about this before. What to use. Is an interesting concept, because if you come to these spaces every week and you're like, Oh, this model is now state of the art, that model is state of the art.
[00:50:05] You may end up not building anything, because you just won't have the, you always keep chasing the latest and greatest. The differences are not vast from week to week, we're just seeing like better scores. But it's well worth checking out this effort for the methodology, for the confirmation that you have.
[00:50:21] Let's say you, you felt that Mistral is better and now you can actually understand. And also for friends of the pod I think John Durbin is also, Error Boris model is really great and it's also up there. And what Nistan highlighted is that bigger models sometimes excel at different things some summarization or just more knowledge.
[00:50:38] It's also outlined there as well. And You can also see models that are not that great, that maybe look good on the leaderboards, but don't necessarily perform as well, and you can see them as well in that effort.
[00:50:49] So maybe actually let me reset the space. Everybody who joined in the middle of me speaking is like, why is this guy speaking?
[00:50:56] And what's going on here? You are welcome. You, you're in the space of ThursdayAI. ThursdayAI we are meeting every week to talk about everything that happens in the world of AI. If you're listening to this and you're enjoying, you're the target audience, but generally we talk about everything from open source LLMs and now embeddings.
[00:51:13] We, we talk about big company APIs. There's not a lot of updates from OpenAI this week. I think they're quiet and they're going to release everything in a week and a half in their dev day. And, and Tropic obviously, and, and... Cloud and Microsoft and Google, like all these things we cover as much as possible.
[00:51:29] We also cover voice and audio. And in that vein, I want to shout out to friends from Gladia and I'll pin there actually, let me just pin this right now. Gladia just released a streaming of Whisper and I've been waiting for something like this to happen. Sometimes for AI engineers, you don't want to host everything yourself. And you want to trust that, the, the WebSocket infrastructure is going to be there when you don't want to build it out. And I'm not getting paid for this.
[00:51:53] This is like my, my personal, if I had to implement like something like the voice interface with ChatGPT, I would not build it myself. I would not trust my own MLOps skills for that. And so for that Gladia is, I've been following them since I wanted to implement some of their stuff and they just implemented like a WebSocket.
[00:52:11] Whisper transcription streaming, and it's multilingual, and it's quite fast, and I definitely recommend folks to check it out. Or check out my review of it, and try out the demo, and if you want it, use it. Because we've talked last week about the interface for ChatGPT that's voice based, and you can actually have a FaceTime call with ChatGPT, and that's incredible.
[00:52:30] And I think more and more removing the screen out of this talking to your AI agents, I think, with the latest releases also in text to speech, like 11 Labs and XTTS that we've covered as well. With advances there, with speed, you can actually start getting interfaces where you can talk, and the AI listens and answers back to you very fast.
[00:52:52] Worth checking out, and definitely an update. Thank you.
[00:52:57] Nisten Tahiraj: Okay. So this is a complete product. I was,
[00:53:00] Alex Volkov: yeah, this is a full, pay a little bit, get a WebSocket and then you use this WebSocket to just like stream and you can embed this into your applications like very fast. Setting that up, I think Koki, you can do this with Koki, which we also covered.
[00:53:13] Gradio Interview with Abubakar, Xenova, Yuichiro
[00:53:13] Alex Volkov: Alright, I think it's time to again, reset the space. ThursdAI I wanna thank Bo who is still on stage. Bo, you're welcome to keep, stay with us a little bit and now we're moving on to the second part of this.
[00:53:30] Welcome, Abubakar. Welcome, Zinova, Joshua. Welcome some folks in the audience from Hugging Face. It's great to see you here on ThursdAI, well, Zinova is always here, or hopefully, but Abubakar, I think this is your first time.
[00:53:41] I'll do a brief intro, and then we can, we can go and talk about Gradio as well.
[00:53:45] I my first inference that I ran on a machine model was a year and something ago, and this was via Gradio, because I, I got this weights file, and I was like, okay, I can, I can probably run something with CLI, but how do I actually visualize this? And back then, Gradio was... was the way and I think since then you already guys you were already part of Hug and Face and Everybody who visited a model page and tried a demo or something probably experienced Gradua even without knowing that this is what is behind all the demos So welcome, please feel free to present yourself.
[00:54:17] Give us Maybe two line, three line of how you explain Gradio to folks, and then we can talk about some exciting stuff that you guys have released this week.
[00:54:25] Abubakar Abid: Awesome. Yeah, first of all, thank you again for, for having me and for having several folks from the Gradio team here. I've known you, Alex, for a long time.
[00:54:32] I think you were one of the early users of Gradio or at least one of the early users of Gradio blocks and, and some of these viral demos. So I've seen, this podcast develop over time and it's It's a real honor to be to be able to come here and to be able to talk about Gradio.
[00:54:45] Yeah. Hi everyone. I'm Abu Bakr. I'm, I lead the Gradio team at Hugging Face. So Gradio is basically the way we describe it is it's the fastest way to, to build a GUI or an app from a machine learning model. So traditionally have, taking a machine learning model to production or at least letting...
[00:55:01] Users try it out has meant that you need to know a lot of front end. You need to know how to, setting up a server, web hosting. You have to figure all of these things out so that other people can play around with your machine learning model. But Gradio lets you do all of that with just a few lines of Python as I think Joshua was mentioning earlier.
[00:55:18] And Gradio has been used by a lot of people. We're very lucky that, we kind of coincide. We started Gradio a few years ago late 2019. It grew out of A project at Stanford, and then spun out to be a startup, and then we got acquired by Hugging Face, and we've been growing Gradio within that kind of ecosystem.
[00:55:32] But we're very lucky because during this time has coincided with a lot of real developments in machine learning. I come from an academic background, so before 2019 I was doing my PhD at Stanford. And, everyone's been doing, machine learning for a while now, but...
[00:55:45] The types of machine learning models that people wanted to build, you built it, you published a paper and that was it, but, since then, recently people are building machine learning models that other people actually want to use, other people want to play around with, things have gone very, exciting, and so that's led to a lot of people building radio demos I think, I was looking at the the stats recently we have something around more than three, four million demos, Gradio demos that have been built since we started the, library.
[00:56:09] And yeah, so recently we released something called Gradio Lite, which lets you run...
[00:56:13] Gradio effects on the open source LLM ecosystem
[00:56:13] Alex Volkov: Wait, before, before, Abubakar, if you don't mind, before Gradio Lite, let's not I just want to highlight how important this is to the ecosystem, right? I'm oriJinally a front end engineer I do component libraries for breakfast, and basically, I don't want to do them it's really nice to have a component library maybe Tailwind UI, or ShadCN, like, all these things, so even front end engineers, they don't like building things from scratch.
[00:56:35] Switching to machine learning folks who like build the model, let's say, and want to run some inference, that's not their cup of tea at all. And, just thinking about like installing some JavaScript packages, like running NPM, like all these things, it's not like where they live at all. And so what Gradio allows us to do this in Python.
[00:56:51] And I think this is, let's start there. That's on its own is incredible and lead, led to so many demos just look to happen in Gradio. And you guys built out pretty much everything else for them, like everything that you would need. And I think recently you've added stuff, before we get to Gradual Light, like components like chat, because you notice that, many people talk to LLMs, they need the chat interface, right?
[00:57:10] There's a bunch of multi modal stuff for video and stuff. Could you talk about, the component approach of how you think about providing tools for people that don't have to be designers?
[00:57:20] Abubakar Abid: Yeah, absolutely. So yeah, that's exactly right. Most of the time when you're, machine learning, developer you don't want to be thinking about writing front end, components that coupled with some, an interesting insight that we had with machine learning models.
[00:57:31] It's much more like the components from machine learning models are tend to be much more usable than in other kinds of applications, right? So one thing I want to be clear is that Gradio is actually not meant to be like a, build web apps in general in Python. That's not our goal. Our goal, we're heavily optimized toward building machine learning apps.
[00:57:50] And what that means is, the types of inputs and outputs that people tend to work with are a little bit more, contained. So we have a library right now of about 30 30, Different types of like inputs and outputs. So what does that mean? So things like images, image editing video inputs and outputs, chatbots as outputs JSON, data frames, various types of inputs and outputs that components that come prepackaged with Gradio.
[00:58:15] And then when you build a Gradio application, you basically say, Hey, this is my function. These are my inputs, and these are my outputs. And then Gradio takes care of everything else, stringing everything together sending them, message back and forth, and pre processing, post processing everything in the right way.
[00:58:29] So yeah, you just have to define your function in the backend, and your inputs and your outputs, and then Gradio spins up a UI for you.
[00:58:36] Alex Volkov: And so I really find it funny and I sent the laughing emoji and said that Gradio was not meant to build web apps, like full scale web apps, because I think the first time that we've talked, you reached out, because I joined whatever open source that was running for stable diffusion, this was before automatic, I think, and you told me hey, Alex, You did some stuff that we didn't mean for you to do, so I injected like a bunch of JavaScript, I injected a bunch of CSS, like I had to I had to go with my full on like front end developer, I was limited with this thing, and so I, even despite the limitation, I think we did like a bunch of stuff with just like raw JavaScript injection, and since then I think it's very interesting, you're mentioning like Gradio demos, Gradio demos, Automatic 1.
[00:59:16] 1. 1, which is maybe for most people, is the only way they know like how to run stable diffusion, is now getting investments from like NVIDIA and getting right, I saw like a bunch of stuff that Automatic does, so it's very interesting like how you started and how the community picked it up. So can you talk about like the bigger parts of this, like Automatic and some other that are like taking Gradio and pushing it to the absolute limit?
[00:59:37] Abubakar Abid: Yeah, absolutely. So that's yeah we're, we're, I'm, I'm, like, perpetually shocked by Automatic 111, every time I see a plug in, or, kind of the, the I think, like you said NVIDIA, now IBM, or something, released a plug in for Automatic 111? It's crazy. But yeah, so basically it's ever since we started Gradio, we've been noticing that, okay, okay, Gradio seems to work for 90 percent of the use cases, but then the last 10 percent people are pushing the limits of, of what's possible with Gradio.
[01:00:06] And so we've progressively increased what's possible. So in the early days of Gradio, there was actually just one class called Interface. And what that did was it allowed you to Specify some inputs and some outputs and a single function. And we quickly realized, okay, people are trying to do a lot more.
[01:00:20] So then about a year and a half ago, we released grad your Blocks, which allow you to like have arbitrary layouts. You can have multiple functions, string together connect inputs and in different ways. And that is what kind of allowed these very, very complex apps like automatic 1 1 1 SSD Next, and the equivalence in other domains as well.
[01:00:36] Of course, like the, the text, the text web, the, the UBA Booga. Text UI as well and then there's also similar kind of, very complex demos in the audio space as well. And music generation as well. So like these super complex, multiple tabs, all of that, that's possible with this new kind of architecture that we laid out called GradioBlox.
[01:00:55] And GradioBlox is this whole system for specifying layouts and and, and functions. And it's defined in a way that's intuitive to. Python developers, the, we like a lot of these like web frameworks in Python have, have popped up. And one of the things that I've noticed as someone who knows Python, but really not much JavaScript is that they're very much coming in from the perspective of a JavaScript engineer, and so like this kind of React inspired kind of frameworks and, and stuff like that.
[01:01:21] And, and what, that's not very intuitive to a Python developer, in my opinion. And so we've defined this whole thing. Where you can, have these, build these arbitrary web, kind of web apps, but still in this Pythonic way. And we're actually about to take this a step farther, and maybe I can talk about this at some point but next week we're going to release Gradio 4.
[01:01:38] 0, which takes this idea of being able to control what's happening on the page. To the next level. You can have arbitrary control over the ui, ux of any of our components. You can build your own components and use them within a Grady app app and get all of the features that you want in a grad app.
[01:01:52] Like the, the, API usage, pre-processing, post-processing. Everything just works out of out of out of the box. But now with your own kind of level of control, yeah. Awesome.
[01:02:01] Alex Volkov: And it's been honestly great to see just how much enablement. Something like as simple as Gradio for folks who don't necessarily want to install npm and css packages.
[01:02:11] There's not much enablement this gave the open source community because People release, like you said, different significant things. Many of them, maybe you are not even aware of, right? They're running in some Discord, they're running in some Reddit. It's not like you guys follow everything that happens.
[01:02:23] Gradio local URL via Gradio Proxy
[01:02:23] Alex Volkov: Additional thing that I want to just mention that's very important that. When you run Gradio locally, you guys actually expose it via like your server, basically my local machine. And that's been like a blast that that's been like a very, very important feature that people may be sitting behind the proxy or everything.
[01:02:39] You can share your like local instance with some folks, unfortunately only for 72 hours. But actually
[01:02:44] Abubakar Abid: that's about to change. So in 4. 0, one of the things that we're trying to get, so actually, we've been very lucky because Gradio has been developed along with the community. Like you said, like often times we don't know what people are using Gradio for until, they come to us and tell us that this doesn't work, and then they'll link to their repo and it's this super complex Gradio app and we're like, what?
[01:03:01] Okay, why are you even trying that? That's way too complicated. But, but, then we'll realize like to the extent to what people are building. And so this you mentioned the share, these share links as well, which I want to just briefly touch upon. So one, one of the things that we released in like the early days of, of, of Gradio is we realize People don't want to worry about hosting their machine learning apps.
[01:03:19] Oftentimes you want to share your machine learning app with your colleague. Let's say you're like the engineer and you have a colleague who's a PM or something who wants to try it out. Or it might be if you're in academia, you want to share it with fellow researchers or your professors, whatever it may be.
[01:03:33] And like, why do all of this hosting stuff if you just are, are, like, building an MVP, right? So we built this idea of a share link. So you just, when you launch your Gradio app, you just say share equals true. And what that does is it creates a it uses something called Fast Reverse Proxy to actually expose your local port to a to this FRP server which is running in a public...
[01:03:53] Machine, and what that does is it forwards any request from a public URL to your local, port. And, what the, in a, the long story short, what that does is it makes your Gradio app available on the web for anyone to try. It runs for 72 hours by default, but now what we're doing as part of 4.
[01:04:08] 0, we'll, announce this, is you can actually build your own share servers. So we have instructions for how to do that very easily and you can point your Gradio instance to that share server. So if you have an EC2 instance running somewhere, just point to it and then you can have that share link running for as long as you want and you can, share your share servers with other people at your company or your organization or whatever it may be and they can use that share link and, again, they can run for however they want.
[01:04:30] Wait,
[01:04:31] Nisten Tahiraj: wait, wait, is this out? Which branch is this? Is
[01:04:34] Abubakar Abid: this going to be out? This is going to be out on Tuesday for Gradio 4. 0 we're, we're going to launch on Tuesday.
[01:04:41] Nisten Tahiraj: It's like the most useful feature I'd say of, of Gradio, especially when you make a Google collab that you want people to just run in one click.
[01:04:49] And like, how are they going to even use this model? And you just throw the entire Gradio interface in there and you share equals true. And then they know, they can just give it, give the link to their friends and stuff. It's really, it makes it really easy, especially with Google Colab. But now that you can host your own, this is huge.
[01:05:09] This is going to... to another level. I have more questions for
[01:05:14] Alex Volkov: Google. I think I just Nissen, thank you. I just want to touch upon the Google collab thing. I think at some point Google started restricting how long you can run like a collab for, and I think you guys are the reason. This exact thing that Nissen said.
[01:05:30] People just kept running the Gradio thing with the URL within the Google collab and exposing like stable diffusion. They didn't build collab for that, and I think they quickly had to figure out how to go around this.
[01:05:41] Abubakar Abid: Yeah. And their approach is like literally blacklisting the name of the the of specific, GitHub repos, which, I, I completely understand where, where Colab is coming from, right?
[01:05:50] They're giving these GPUs for free. They have to have to prioritize certain use cases, but we're working with the Colab team and we're seeing if, there's ways, like right now it's like a blacklist on, on automatic one on one, and some other repos. So we're hoping we can find another way that's not That's not so restrictive.
[01:06:05] Nisten Tahiraj: No, but it still works. You can just fork the repo. It works for everything else. It works for LLMs. So if anybody else really needs it. Gradio works on Colab. Well, as far as language stuff goes, I haven't done that much.
[01:06:18] Abubakar Abid: Yeah, so Gradio works on Colab for sure. And, and that's, and that's early on, like one of the decisions we had to make actually was...
[01:06:25] Should we use like, the default python runtime or should we like change, like the interpreter and stuff like that? Because building GUIs is not necessarily python's like strength, and like oftentimes you wanna render re-render everything, and you, you wanna do certain things that may not be like what Python is suited for.
[01:06:42] But early on we decided, yeah, we wanna stick with the default python runtime because. One of the reasons was things like Colab, because we wanted people to be able to run Gradio wherever they normally run Python without having to change their workflows. And Colab, Gradio works in Colab.
[01:06:56] We had to do a lot of... Trickery to make it work. But yeah, it works. It's just like these certain very, very specific apps that have become too popular and apparently consume too many resources. They're blacklisted by Colab right now.
[01:07:10] Local inference on device with Gradio - Lite
[01:07:10] Alex Volkov: Alright thank you for this intro for Gradio. To continue this, we have on stage Zinova who introduced himself, authors of TransformerJS, we've been talking with Bo in the audience, also somebody who's like just recently open sourced, with Jina, the embeddings model, and everything that we love to cover in ThursdAI, a lot of it is talking about As open source, as local as possible, for different reasons, for, not getting restricted reasons.
[01:07:36] And you guys just recently launched Gradio Lite, and actually we have Yuichiro here on stage as well. So I would love to have you, Abubakar introduce and maybe have Yuichiro then follow up with some of the stuff about what is Gradio Lite? How does it relate to running models on, on device and open source?
[01:07:52] And yeah, please, please introduce it.
[01:07:54] Abubakar Abid: Yeah, absolutely. Yeah. Like you mentioned, I think one of the things that we think about a lot about at Gradia is like, it's the open source ecosystem and, right now, for example, where can open source LMs, for example, really shine and things like that.
[01:08:06] And one of the places is on device, right? On device or in browser is, open source has a huge, edge over proprietary models. And so we were thinking about how can Gradio be useful in this setting. And we were thinking about the in browser application in particular. And we were very, very lucky to have Yuchi actually reach out to us.
[01:08:25] And Yuchi has this, fantastic tracker, but if you don't already don't know Yuchi, he built Streamlit Lite, which is a way to run Streamlit apps in the browser. And then he reached out to us and basically had this idea of doing something similar with Gradio as well. And basically, I, almost like single handedly refactored much of the Gradio library so that it could run.
[01:08:43] In, with Pyodide, in WebAssembly, and basically just run in the browser. I'll let Yuchi talk more about that, but basically, if you know how to use Gradio, then you know how to use Gradiolite. You just write the same Python code, but wrap it inside Gradiolite tags, and then it runs within the browser, like in the front end.
[01:08:59] You can execute arbitrary Python, and it just works. Yuchi, if you want to share a little bit more. About that, yeah, or introduce yourself.
[01:09:08] Yuichiro: All right, Hey, can you hear me? Well, thank you very much, thank you very much for the quick short interaction about Gradual Light and Streaming Light 2.
[01:09:15] Well, as Abhakal explained about it,
[01:09:18] there was
[01:09:18] Sorry. OriJinally there was kind of a, a tech technological movement about edge computing or Python. It was started by uh ide that was c Python runtime compiled for web assembly that can completely run on web browsers. It started, it triggers the big band of edge computational Python, random starting with project that was.
[01:09:43] That was already deported to Web Assembly runtime as D Light, and it inspired many other Python frameworks, including Streamlet and any other existing Python frameworks. I dunno uh pet script or. HoloVis panel or Shiny for Python, something like that. So there was a huge movement about to, to make Python frameworks to compatible with WebAssembly or web browser environment.
[01:10:13] And I thought that was a great, opportunity to make machine learning or data science stuff completely run on web browser, including transformer things or many more stuff existing in the stream machine learning ecosystem. And I first created a Streamlit Lite that was forked version of Streamlit to WebAssembly.
[01:10:36] And yeah, the, remaining story were the same as what Abhakaal introduced. So technically it was not, my oriJinal stuff, but there was a huge movement about such kind of stuff. And I, simply followed that. flow and my the transfer such kind of analogies to gradial repository.
[01:10:58] Alex Volkov: Yeah, that's it. Awesome. Thank you so much. And okay. So can we talk about what actually do we do with now the ability to run Gradio all in the browser? Could maybe both of you give some examples and then I would like also to to add Zenova to the conversation because much of the stuff is using Transformers.
[01:11:18] js, correct? Can we maybe go and talk about what is now actually possible compared to like when I run Gradio on my machine with a GPU and I can run like Stable Diffusion? I just
[01:11:27] Nisten Tahiraj: want to say that this is crazy that this can happen at all for the audience to
[01:11:32] Abubakar Abid: prepare. Yeah, I was honestly blown away the first time Yuchi showed me a demo as well.
[01:11:36] Imagine you have a, any sort of machine learning model. Practically, not almost anything, but a super, really good speech recognition model running completely in your browser. Meaning that, for example, now you can take that demo, you can put it inside GitHub Pages.
[01:11:51] You can host it inside. We've seen people embed Gradio demos now with Gradio Lite inside Notion. So you have a Notion whatever page, you can take that demo, you can embed it inside Notion. One of the things that we launched when we launched Gradio Lite at the same time is we also launched something called the Gradio Playground.
[01:12:07] Now the Gradio Playground, you can actually just Google this, you can find this. But basically what it allows you to do is it allows you to write code in the browser. And as you're editing the code, you can see live previews of your Gradio application. And, and basically what's happening is, is taking that Gradio code, it's wrapping it inside Gradio Lite tags, and it's just running it.
[01:12:27] It's just straightforward application of Gradio Lite. And, we're excited by this personally just because if one, it opens up, it allows us to write interactive documentation. You can write, you can try stuff, you can, you can immediately, see the results. We're also excited because we've seen interest from other libraries including, for example, SacketLearn, who want to embed Gradio demos within their documentation.
[01:12:49] Within their docs, right? But they're they were hesitant before because they didn't want to have a separate server running these radio applications and have to worry about maintaining those servers, making sure they were up all the time, making sure they could handle the load. Now they can write it in their docs and they're like, their, their demos and everything, they'll just run in the user's browser.
[01:13:07] They won't have to worry about maintaining everything since it's, in the same code base and everything. So I think that's another kind of cool application that we're excited by is just... These potential for interactive documentations that, maybe potentially other, other maintainers or other libraries might want to include.
[01:13:22] So yeah, so stuff like, security, privacy, serverless type stuff, hosting, and all of that. And then also like these interactive documentations.
[01:13:30] Alex Volkov: I think the demo that you mentioned with the transition within Notion from VB from Hive Interface, I think that was great. I'm trying to find the actual link, but basically, because Notion allows to embed like basically iframes, right?
[01:13:42] So he embedded this whole Gradio Lite interface to translate. I think using... Burt or something like very similar that all runs within like the notion page. I think that's awesome. Joshua, you want to chime in here and say how transformers is built into this and now this allows for way more people to use transformers in like a UI way.
[01:14:02] Transformers.js integration with Gradio-lite
[01:14:02] Xenova: Yeah, sure. So first of all, literally nine, like almost, I would say the whole. Everything that we are talking about now has been, like,
[01:14:12] Abubakar Abid: led by the Gradio team.
[01:14:14] Xenova: And I am here piggybacking and be like, Whoa, look at this, Transformers JS is now working. That's really not what we're talking about today.
[01:14:23] It's the amazing work that the team has been able to do. To achieve the past, this has been going on for, for quite a while. It's been like codenamed like Gradioasm and now finally being released as as GradioLite and now. Sort of like the Transformers J side of it. Just oh, by the way there's this library called Transformers J you can sort of use it and, and with the, the Transformers.
[01:14:48] Oh, was that ? Sorry. You've been way too humble.
[01:14:51] Abubakar Abid: No, no, absolutely
[01:14:52] Xenova: not. I think it's, it's, it's so much has been done by, by you and your, and, and the amazing radio team that's it, it just so happens to be that these things are like coinciding. And now you can end up using Transformers. js with, with Gradio and Gradio Lite.
[01:15:07] And obviously this is also made possible by, okay, everyone, everyone stick with me. It's going to be a little, get a little complicated when I try to explain this. But Transformers. js Pi, which is, are you ready? A JavaScript port of a Python library turned into a JavaScript library so that I can run in a Python environment.
[01:15:29] Okay. We all caught up? That's, that's Transformers. js. py, which is which, which Yushi wrote in the audience obviously with his experience with streamlets bringing streamlets to the browser. It's sort of his, his invention, which is quite funny, but that's sort of how Transformers. js is able to be run.
[01:15:49] Inside Gradiolites there are other ways, but from what you'll see in the documentation, that's sort of like the, the go to way. And it's
[01:15:57] Alex Volkov: Yeah, I wanna ask about this, because I saw from Transformers. js, import import on Discord Transformers. js.
[01:16:04] So maybe you should, could you talk about this part that, that Zenova tried to explain? It was, like, a little A little complex Transformer. js is you can install it through npm and then run this, right? And then it runs in the, in the node environment and browser environment. Gradualite is basically Python within JavaScript.
[01:16:19] So then you have to turn Transformers into Python in order to get it into Gradualite so that it runs within the JavaScript context again? Is that, is that correct? Am I getting this right?
[01:16:30] Nisten Tahiraj: If I could say something for the audience, what has, what's happening here is that there's a layer called Pyodide and that uses kind of like what WebAssembly uses to run Python at native speeds.
[01:16:44] So it runs in the browser. And it goes down that stack, there's like a virtual machine and compiler, all that stuff in there. And then that's how Python is able to run at native speed. And this means that with PyScript, you can have inside the same index. html, just your regular index. html, you can have your JavaScript code and your objects and stuff.
[01:17:06] And you can have just straight Python code in there. Like you just add the tag. You just dump the Python as is nothing else. And the crazy part is that it can access JavaScript objects now. So you can do the math in Python in the browser, 'cause JavaScript can do math well again, but, and then you can access those objects.
[01:17:30] So this is a whole crazy stack here with PIO IDE and EM scripting. And again, that's only WebAssembly. So that's CPU. Only for now, because there's still a mountain of work to get, and to finish it off, Emscripten is like your POSIX layer, like your Unix layer. It's like there's an operating system being built inside the browser here that's going on.
[01:17:54] So that's why things are getting complicated, but yeah, just to keep that in mind, that's the base.
[01:17:59] Yuichiro: Yeah, but, what Nisten talked about was everything, because we can access the JS object from Python world inside the browser if you import transformer.
[01:18:10] js. py on Gradle Lite, under the hood, transformer. js is now still being imported in the browser environment. And what... When you write a Python code as a Gradle like application on the browser, what you do is, simply using the oriJinal JavaScript, JavaScript version of Transformer. js, just proxy from the Python code through the, proxying mechanism provided by Pyodite.
[01:18:42] What Transformer. js. py does is just a, thin, Proxying layer or some glue code between bridging these two, two words, Python and JavaScript. That's it.
[01:18:56] Abubakar Abid: Yeah, I just zooming out a little bit. So basically what, what Transformers. js underscore pi does, it lets you run everything that Transformers.
[01:19:03] js does. And what Transformers. js does, it lets you run a lot of the models, a lot of the tasks. There's a lot of the models there, you can now run in your browser, right? We're talking about all of the NLP related tasks, like things like translation, LLMs, but also, a lot of the vision tasks, a lot of the audio stuff.
[01:19:22] We're talking about speech recognition that's powered by Transformers, what Josh has been doing with Transformers. js. And I think Transformers. js just released, even for example, speech generation. Text to speech. And so now you can do that within Transformer. js, which means you can do it within Transformer.
[01:19:34] js Pi, which means now you can do it within Gradial Light as well.
[01:19:40] Alex Volkov: That's incredible. And I think the biggest part for me is that... Now that you guys ported Gradio, which is ubiquitous in machine learning, and everybody who releases the model uses either this or Streamlit, but I think it's, it's a clear winner between the two, as this is, as I'm concerned and as I see, then now you basically ported the same thing towards the browser, and the more we see, Models getting smaller and we've been always talking about this models getting smaller, models being uploaded to the browser.
[01:20:08] Browsers
[01:20:09] Abubakar Abid: getting more powerful and,
[01:20:10] Alex Volkov: and WebP is more browser getting more powerful. Yeah. And yeah, I'm getting, I'm getting to web GPU because we have backend here, Arthur, on stage. And I would love to introduce you guys if, unless you're already familiar. The more we see this move, the more like the need for something like a component library that's built in is, is very interesting.
[01:20:25] Even though this world already has a bunch of libraries. But you're basically, with this, you're also porting the people with the experience of Gradio, right? With the existing with the existing frameworks, with the existing Gradio interfaces, to this world. I find it very exciting, so thank you.
[01:20:38] And I want to introduce Arthur. Arthur, feel free to unmute yourself and maybe introduce yourself, briefly, and then yeah, feel free to chime in to this conversation.
[01:20:46] Arthur Islamov: Okay, so I did quite a lot of things with ONNX to create the Diffuser. js library and to load stable diffusion in the browser, and now I'm working on the SDXL version.
[01:20:58] So I was going to ask, do you know if there are some plans on adding WebGPU backends for PyTorch? Because when it happens... It'll be so much easier as web GPU backend can be launched on any platform, not even in the browser, but also locally without the browser, just using the metal backend the direct ticks or Vulcan or Linux.
[01:21:28] So I guess when that happens, we'll go to a whole new era as you'll be able to run those PyTorch models in the browser with GPU acceleration.
[01:21:40] Xenova: I can tag on to this. The TLDR of it is it's not at the point...
[01:21:46] Where, sort of, I'm sort of comfortable with upgrading the ONNX Runtime web Runtime, basically, to support the WebGPU backend right now, just because there's quite a few issues still left, like left, so we'll see. To solve before we get to the point where you can start running these models completely on WebGPU.
[01:22:07] The main, I think, the current issue at the moment is with, like, when you're generating text a lot of the... The buffers aren't reused properly during when you, when you start decoding. That's sort of leading to quite a massive performance bottleneck just because you're transferring memory between CPU and GPU every single time you're you're decoding.
[01:22:31] So that's, that's not quite there yet, however, with things like image image classification and I guess models with encode only, encoder only models, those are getting quite good, like births pretty fast we've, segment anything when you're just doing the encoding step, we the Onyx Runtime team has got to the point where it used to take around 40 seconds and now it takes around 4 seconds.
[01:22:55] And that's currently being worked on in like a dev branch, basically, of Transformers. js, just like making sure the integration's working. But it's, it's almost there. I keep, I keep saying it's almost there, but the amazing Microsoft team has been, has been really working hard on this. And if you just look at the commit history of on GitHub Microsoft slash Onyx Runtime and you go to the web version.
[01:23:18] There's just so many amazing people working on it and it's slowly getting to a point where and this will sort of be released with Transformers. js version 3. When we upgrade the Onyx runtime version to probably 1. 17, which will be, which will be the next one. It's currently 1. 16. 1. And then they'll, it's, and, and literally from the user's perspective, it's as simple as adding a line of code just saying, Basically use web GPU instead of web assembly.
[01:23:46] And, and also in the case where it's not supported, it'll fall back to the web assembly implementation. And, and that's, this will completely be transferable to how grid your light works, just because as was mentioned, it's sort of use as transformers js under the hood. So you any benefits that you'll see in Transformers j you'll see in transformers Jss pie, which you'll see in radio lights, which is, which is great.
[01:24:11] TLDR coming soon, it's an, it's an annoying answer to give, but it's, it's so close. And I guess this is also good because it sort of aligns with the time that more browsers will support WebGPU, sort of like without flags. I know Chrome is sort of leading the charge and other Chromium based browsers.
[01:24:30] But if you look at things like Safari and Firefox, it's quite far behind to the point that you it's. It's not, it's not ready for like mass adoption yet, but once it is, and once the ONNX Runtime backend has the WebGPU support has improved, you'll definitely be seeing that in Transformers Jest.
[01:24:48] So hopefully that answers the question. I
[01:24:52] Nisten Tahiraj: think stuff's about to get crazy on the front end because of this because the thing about you have all your WebGL stuff, you have all your maps, all your 3D, all your games. Now you can have an LLM even generate code for them, manipulate those objects, move stuff around on screen in 3D, and like the, the AI.
[01:25:14] Does that at all, all within your machine, but I do want to say that for, for Pyodide itself, it might take a long time for a WebGPU support because it depends on EM scripting. And if you want to do anything with Python, like open a file, write a file, output a file, you only can do what EM scripting gives you and EM scripting is like the base layer.
[01:25:39] Of the operating system, like it pretends, it fools your apps into thinking that there's an operating system there when, when there isn't. And as far as I've seen, like two, three months ago, WebGPU support was like really, really early on and might take a while for Emscripten to support that. So you're going to have to do that other ways by going straight to using WebGPU versus using it with that layer.
[01:26:06] So it might get a bit complex
[01:26:09] Alex Volkov: there. I agree about the stuff is about to get crazy. Go ahead, Arthur, and then we'll follow up on Gradio 4 and then we'll conclude.
[01:26:18] Nisten Tahiraj: Yeah, I
[01:26:18] Arthur Islamov: just wanted to tell that as yesterday or a few days ago I have seen that This distilled stable diffusion model I saw that they have previously released not Excel version, but the ordinary 2.1 or something like that, the distilled one.
[01:26:35] So I'm thinking to try to make my edema work with that distilled model without 64 beats. So for just ordinary 32 bit that will work in almost any browser without any additional flags or launching with some special parameters.
[01:26:54] Alex Volkov: Yeah. Arthur, you can't just mention an item on my updates list and not talk about this, right?
[01:26:59] Folks, let me just briefly cover what Arthur just said. Just on the fly SegMind, a company called SegMind, introduced like a distilled version of SDXL. And it's okay, Diffusion Excel, something they released a while ago. We've covered this multiple times. Understand like way better. Quality, obviously, generations and diffusion, but also way better text understanding, right?
[01:27:20] And it has two parts there's like a refiner part in addition. And so this company basically distilled that distillation we've talked about multiple times before. It's when you train your own model, but then you steal data from GPT 4 and you create the data so that GPT 4, you basically distill, it's like smartness to your own models.
[01:27:37] So they basically did this for SDXL they call it SegMind Stable Diffusion 1B, and it's a. 50 percent smaller and 60 percent faster than SDXL. Again, just to put in some time frames what Abubakar and I talked about, where I first experienced Gladio, this was Stable Diffusion 1. 4, a year ago, a year and a couple of months ago.
[01:27:57] Since then, we got Stable Diffusion multiple iterations of Stable Diffusion, then there is SDXL, which is like the Excel version. It it generates 124 by 124 images. And, and then now a few months after they released that, now we have a version that's 50 percent smaller and 60 percent faster.
[01:28:16] And so what Arthur is like now talking about diffusers JS is the ability to like load stable diffusion in the browser. Now there's a model that's half the size, and 60 percent has passed, which is good for the browser context. So I pinned it to the top of the tweet check out SegMind, it's definitely super cool.
[01:28:34] And the advancements that we see from week to week, and this is obviously super cool as well. And Arthur, sorry to interrupt with this, but you had one of my tasks that I had to finish before we finish and talk about this. So are you, have you already introduced it to diffusers? Have you tried it?
[01:28:52] Nisten Tahiraj: I have
[01:28:52] Arthur Islamov: tried it to convert it to omx, but it didn't work or maybe some of my code didn't work. So I guess I will try again on the weekend and yeah, most likely I will make it running.
[01:29:06] Alex Volkov: I think we had some folks on segment react and we will, let's try to connect there and, and hopefully get it running on, on as well so that we all benefit and I guess, maybe as the last part of this conversation Abukar and thank you for joining Uchi ABA Ali and the folks from Hugging Face. It is great to see all of you. I think you mentioned some folks that joined before like Uhb and some other folks on the Hugging Face. We're big fans here on ThursdAI, and we're like, always welcome you guys.
[01:29:33] Could you talk about what's coming in version four, because I think you, you, you gave us like one tidbit. But give us an update on that. I would love to hear.
[01:29:40] Abubakar Abid: Yeah, yeah, definitely. Yeah, so we're launching Gradio 4. 0 on Tuesday October 31st. And basically, the, the team has been working very, you mentioned earlier that, people are, are building these very, very complex apps with, with Gradio and, and really, honestly, stuff that we did not anticipate when we were designing Gradio.
[01:29:57] And more and more, what we want to do is almost take ourselves out of this feed, feedback loop. And let people build what they want to build, but let the community build stuff that, whatever you imagine, kind of just be, just be able to put that in a Gradio app. Let me be a little bit more concrete.
[01:30:11] So what is Gradio 4. 0 going to introduce? For example it's going to introduce the idea of custom components. So if you know a little bit of Python, a little bit of JavaScript, you can build your own component. You can use that within a Gradio app, just like you do normally, just like you use our built in, 30 or so built in components.
[01:30:27] Speaking of the built in components, we're redesigning some of the components from scratch, particularly the media components. So things like image audio video, they're going to be much, much nicer and they're going to be fully accessible. So one of the things that we're realizing is that, we're not, at Gradio, we're not just building a product for a specific audience, but we're building tools that let people build, apps for many different audiences.
[01:30:50] And so we want to make sure that all of the core components are accessible. That way it's easy to do the right thing and build accessible web applications. So we're redesigning that we're switching over from WebSockets to as server side events. There's several reasons for this and, and we'll talk about more about this on, on Tuesday.
[01:31:07] We're, we're having a little long live stream as well, but there's several reasons why server side events is the way to go for Gradio. And so there's, that's more of an internal refactor. You probably won't notice things. You might notice some speed ups in certain situations. It'll unlock a lot of things later on.
[01:31:22] We're open sourcing the the the sharing links process, the share servers at Gradio. So everyone will be able to set up their own, custom share links. So instead of, whatever dot Gradio dot live, you can have, you can have, some, some code, dot turjom dot video if you want, you can have whatever URL, custom URL you want for your share links.
[01:31:42] And then a lot of other changes as well we'll, we'll, we'll talk more about that on Tuesday. The team has been working super hard, so I'm, I'm excited to, to get it out for you guys to try out.
[01:31:51] Alex Volkov: That's so awesome, and, and can't, can't wait to see this. I I think the share links is like such a powerful virality thing, that once people start adding this to their domains, and, and start running different Gradio interfaces within Colab, outside of Colab, with their own domains.
[01:32:08] I think it's going to be super cool, especially if they don't expire. I absolutely received many of these links over DMs from multiple people. I think even people in the audience so far. And I think adding them to the custom domains. Thank you for open sourcing that. That's great.
[01:32:21] Abubakar Abid: I think part of it is also we want to reduce the load on our shared servers.
[01:32:25] We're getting too many of these links being created
[01:32:27] and stuff.
[01:32:27] Alex Volkov: Yes, absolutely. And I think the accessibility features are great. Folks, definitely check out follow Abu Bakr, follow Yuchi, and folks on stage, do Ali as well. To stay tuned to what's coming up to Gradio and then make sure to update your Gradio interfaces to the new accessible ones because you're building it's no longer demos.
[01:32:46] Everybody's using every new model is, is getting a Gradio interface and accessibility is very important for the level of web applications. I think... With that, I want to thank you guys for coming up and sharing with us Radio Light, which is very, very in accordance to what we love to talk about here, open source, open source LLM, on device inference, and taking control of your own, LLMs.
[01:33:07] I think, Nistan, you briefly, briefly talked about how crazy it's going to be. Where there is an LLM built into your website or web application that runs on the GPU of your device and is able to do stuff. And you can interact with it without basically offline. That's great. I think Nishtan, something that I want to talk about, but maybe I'll let you talk about this.
[01:33:27] I will say this thing. Now that we've concluded the interview with Gladio folks, one of the things that we love. Most of all, on Thursday, I have breaking news, and we actually have some breaking news. Nissen, go ahead, please, present the breaking news that you just sent.
[01:33:39] Nisten Tahiraj: I pasted a Gradius space above.
[01:33:43] If you click on it, that's what it is. And it's it's Kokui's, new release, new voice model. This is huge because they're allowing fine tuning on their voice model. And one criticism of the open source voice models has been that the dataset for training them has been of poor quality, like the, the microphone and, and stuff and the, the dataset that people use to, to train the models has been bad.
[01:34:12] So this is. Pretty important in this regard, because it's one of the very few, there's the one that Zenova released, and the Kokui one, that are open source and usable when it comes to text to speech, and that are like somewhat, somewhat pleasant, and that run relatively fast. Otherwise, it's pretty hard to have text to speech.
[01:34:37] Yeah, the, the part that you can fine tune, they, they open source the fine tuning code. Yeah, go there and
[01:34:43] Alex Volkov: get that, yeah. Thank you, Nissan. The, the folks from Cochlea, when they released XTTS, which is the open source text to speech that kind of, we know 11 Labs, we know Play. ht, we know OpenAI has one that Spotify uses to translation, and OpenAI haven't released any.
[01:34:59] We'll see next week if they're gonna give us an API for that. All of those require... A lot of money, just a lot of money 11 labs is basically rolling in cash because everybody wants to get their AIs to talk, right? And so we previously here, we talked about the listen part, we've talked about the streaming from Gladiator, but now you can have Whisper basically streaming.
[01:35:18] The other part of that was, hey, well, once your LLM listens and thinks, which is the inference part, you also want it. to talk to you. And TTS Texas speech is the way to do that. And Kokui, we had a chat with Joshua when they released TTS, which was very exciting. And now live on stage, live on Thursday.
[01:35:34] I, because this is why Thursday exists. Many people release stuff on Thursday. There is their own fine tuning with minutes of data. So you can create a voice, let's say maybe this is going to be disappointing for folks here on stage, but everybody here who spoke on stage more than a minute now is basically public for everybody else to take your voice and clone it with XTDS.
[01:35:56] It was possible before, somebody just had to pay money for it, but now... And Ali's laughing because Ali didn't talk yet, but everybody's basically now is, is going to get voice cloned. It's very easy. We're going towards this future. If this future scares you, there's no escape from that. Even I think VAL E from Microsoft, when it was released, they talked about like maybe 30 seconds of voice is enough to clone.
[01:36:18] But XTTS now gives us basically a framework and, and even the new language they said to, to add to Kokui, to the XTTS, and then you can use this. Within Transformers. Zinova, can we use Cochlear within Transformers. js or not yet? I think, I think we can. Not yet. Not yet. Okay, so soon, soon you'll be able to even do this all completely all within the browser.
[01:36:41] Hopefully once integration with WebGPU lands. So here we have it folks. We had an incredible Thursday I today. We started with talking with Bo and the guys from the embeddings team that released like the, the most kind of up how should I say, most comparable to OpenAI embeddings in open source.
[01:36:59] And that was great. Bo actually gave us like a masterclass in how embeddings work and Jina embedding models are available within, talked with Bakar and Yuchi and ANOVA and Arthur and all on stage the team behind Grado that if you haven't used Grado, you probably have used Grado. You just didn't know that it's Grado And, actually, this interface or slash library that started for demos only scaled all the way up to something like automatic, where like multiple people compute, contribute thousands of contributions including like NVIDIA and and I think IBM contribute now. It's like a full businesses run on this quote unquote component library.
[01:37:37] And I I just want to invite you to join Thursday and next week as well because... Some of this was planned, but definitely not all of this. But also, this is the way to stay up to date. And next week we're going to see some more incredible things. I think some very interesting things are coming up.
[01:37:52] I will have a personal announcement to make that's going to be very surprising to some folks here on stage. But definitely we'll keep ThursdAI going significantly, significantly more. And with that, I just want to thank you for joining us. It's been a pleasure to have these. It's been a pleasure to like have a space where, the graduate folks and the Jina folks can come and talk about what they released.
[01:38:12] And we can actually ask them questions. I want to thank everybody who joined on stage. Nistan, thank you always for joining, Zinova, Arthur. We, we were joined by new folks that we'll introduce next. Thank you so much for joining us this week because we just don't have the time and obviously thank you for folks in the audience who join every week and I see Enrico in there and Junaid and Tony and some other folks that like I love to see from week to week.
[01:38:33] If you missed any part of this, any part at all, or if the internet connection for you got stuck ThursdAI is about a live recording, but then it's getting released as a podcast episode, so you will get, if you're subscribed, and you should be already subscribed to ThursdAI on Apple or Spotify, you'll get this episode hopefully very quickly edited, if I'm not getting lost in some other interesting stuff, like Kokui.
[01:38:57] Thank you. And we will also release a newsletter with all the links and the conversations with Guardia team and, and Bo and all the updates as well in the form of links. And with that, I thank you for joining. It's been two hours. It's been a lovely time and now I need to go and actually edit the podcast.
[01:39:12] See you here next week. Thank you and yeah, please share with your friends as much as possible. The more crowd there is, the better these will be. And yeah, help and participate. Thank you all and have a good rest of your week. Bye bye.
Hey friends, welcome to ThursdAI Oct - 19. Here’s everything we covered + a little deep dive after the TL;DR for those who like extra credit.
ThursdAI - If you like staying up to date, join our community
Also, here’s the reason why the newsletter is a bit delayed today, I played with Riffusion to try and get a cool song for ThursdAI 😂
ThursdAI October 19th
TL;DR of all topics covered:
* Open Source MLLMs
* Adept open sources Fuyu 8B - multi modal trained on understanding charts and UI (Announcement, Hugging face, Demo)
* Teknium releases Open Hermes 2 on Mistral 7B (Announcement, Model)
* NEFTune - a "one simple trick" to get higher quality finetunes by adding noise (Thread, Github)
* Mistral is on fire, most fine-tunes are on top of Mistral now
* Big CO LLMs + APIs
* Inflection Pi got internet access & New therapy mode (Announcement)
* Mojo 🔥 is working on Apple silicon Macs and has LLaMa.cpp level performance (Announcement, Performance thread)
* Anthropic Claude.ai is rolled out to additional 95 countries (Announcement)
* Baidu AI announcements - ERNIE 4, multimodal foundational model, integrated with many applications (Announcement, Thread)
* Vision
* Meta is decoding brain activity in near real time using non intrusive MEG (Announcement, Blog, Paper)
* Baidu YunYiduo drive - Can use text prompts to extract precise frames from video, and summarize videos, transcribe and add subtitles. (Announcement)
* Voice & Audio
* Near real time voice generation with play.ht - under 300ms (Announcement)
* I'm having a lot of fun with Airpods + chatGPT voice (X)
* Riffusion - generate short songs with sound and singing (Riffusion, X)
* AI Art & Diffusion
* Adobe releases Firefly 2 - lifelike and realistic images, generative match, prompt remix and prompt suggestions (X, Firefly)DALL-E 3 is now available to all chatGPT Plus uses (Announcement, Research paper!)
* Tools
* LMStudio - a great and easy way to download models and run on M1 straight on your mac (Download)
* Other
* ThursdAI is adhering to the techno-optimist manifesto by Pmarca (Link)
Open source mLLMs
Welcome to multimodal future with Fuyu 8B from Adept
We've seen and covered many multi-modal models before, and in fact, most of them will start being multimodal, so get ready to say "MLLMs" or... we come up with something better.
Most of them so far have been pretty heavy, IDEFICS was 80B parameters etc'
This week we received a new, 8B multi modal with great OCR abilities from Adept, the same guys who gave us Persimmon 8B a few weeks ago, in fact, Fuyu is a type of persimmon tree (we see you Adept!)
In the podcast I talked about having 2 separate benchmarks for myself, one for chatGPT or any MultiModal coming from huge companies, and another for open source/tiny models. Given that Fuyu is a tiny model, it's quite impressive! It's OCR capabilities are impressive, and the QA is really on point (as well as captioning)
An interesting thing about FuYu architecture is, because it doesn't use the traditional vision encoders, it can scale to arbitrary image sizes and resolutions, and is really fast (large image responses under 100ms)
Additionally, during the release of Fuyu, Arushi from Adept authored a thread about visualQA evaluation datasets are, which... they really are bad, and I hope we get better ones!
NEFTune - 1 weird trick of adding noise to embeddings makes models better (announcement thread)
If you guys remember, a "this one weird trick" was discovered by KaiokenDev back in June, to extend the context window of LLaMa models, which then turned into RoPE scaling and YaRN scaling (which we covered in a special episode with the authors)
Well, now we have a similar "1 weird trick" that by just adding some noise to embeddings at training time, the model performance can grow by up to 25%!
The results very per dataset of course, however, considering how easy it is to try, literally:
It's as simple as doing this in your forward pass
We should be happy that the "free lunch" tricks like this exist.
Notably, we had a great guest, Wing Lian the maintainer of Axolotl, a very popular tool to streamline fine-tuning, chime in and say that in his tests, and among the discord folks, they couldn't reproduce some of these claims (as they are adding everything that's super cool and beneficial for finetuners to their library) so it remains to be seen how far this "trick" scales, and what else needed to be done here.
Similarly, back when the context extend trick was discovered, there was a lot of debates about it's effectiveness from Ofir Press (author of ALiBi, another context scaling methond) and futher iterations of the trick made into a paper and a robust method, so this develompment is indeed exciting!
Mojo 🔥 now supports Apple silicon Macs and has LLaMa.cpp level performance!
I've been waiting for this day! We've covered Mojo from Modular a couple of times and it seems that the promise behind it starts to materialize. Modular promises an incredible unbelieavable 68,000X boost over vanilla python, and it's been great to see that develop.
Today (October 19) they have released their support of Mojo Lang on Apple silicon which most developers use, and it's a native one and you can use it right now via CLI.
A friend of the pod Aydyn Tairov, hopped on the live recording and talked to use about his LLama.🔥 project (Github) that he ported to the Apple silicon, and showed an incredible, LLaMa.cpp like performance, without crazy optimizations!
Aydyn collected many LLaMa implementations, including Llama.cpp, LLama.c by Karpathy and many others, and included his LLama.mojo (or Llama.🔥) and saw that the mojo one is coming very very close to LLama.cpp and significantly beats Rust and Go and Julia examples (on specific baby llama models)
The Mojo future is bright, and we'll keep updating with more, but for now, go play with it!
Meta is doing near-real time brain → image research! 🤯
We've talked about fMRI signals (and EEG) signals being translated to diffusion imagery before, and this week, Meta has shown that while fMRI signals to brain imagery is pretty crazy on it's own, using something called MEG (non invasive Magnetoencephalography) they can generate and keep generating images based on the brain signals, in near real time!
[TK video here]
I don't have a LOT to say about this topic, besides the fact that as an Aphant (I have Aphantasia) I can't wait to try this on myself and see what my brain actually "sees"
Baidu announces ERNIE and a bunch AI native products including maps, drive, autonomous ride hailing and more.
Baidu has just wrapped up their biggest conference of the year, BaiduWorld, where they announced a new version of their foundational model called ERNIE4, which is a multimodal (of unknown size) and is now integrated into quite a few of their products, many of which are re-imagined with AI.
A few examples beyond a basic LLM chat like interface are, a new revamped map experience with an AI assistant (with voice) built in to help you navigate and find locations, a new office management app that handles appointments and time slots called InfoFlow, and it apparently even does travel booking, to an AI "google drive" like product called YunYidou, that is able to find video content, based on what was said and when, and even pinpoint specific frames, summarize and do a bunch fo other incredible AI stuff, here's a translated video of someone interacting with YunYinou and asking for a bunch of stuff one after another.
Disclosure: I don't know if the video is edited or in real time.
Voice & Audio
Real time voice for agents is almost here, chatGPT voice mode is powerful
I've spent maybe 2 hours this week, with chatGPT in my ear, using the new voice mode + AirPods. It's almost like... being on a call with chatGPT. I started talking to it in the store, asking for different produce to buy for a recipe, then drove home and ask it to "prepare" me for the task (I don't usually cook this specific thing) and then during my cooking, I kept talking to it, asking for next steps. With the new IOS the voice mode shows up as a live activity and you can pause it and resume without opening the app:
It was literally present in my world, without me having to watch the screen or type.
It's a completely new paradigm of interactions when you don't have to type anymore, or pick up a screen and read, and it's wonderful!
Play.ht shows off an impressive <300ms voice generation for agents
After spending almost 2 hours talking to chatGPT, I was thinking, why aren't all AI assistants like this, and the answer was, well... generating voice takes time, which takes you out of your "conversation flow"
And then today, play.ht showed off a new update to their API that generates voice in <300ms, and that can be a clone of your voice, with your accent and all. We truly live in unprecedented times.
I can't wait for agents to start talking and seeing what I see (and remember everything I heard, via Tab or Pendant or Pin)
Riffusion is addictive, generate song snippets with life-like lyrics!
We've talked about music gen before, however, Riffusion is a new addition and is now generating short song segments with VOICE! Here are a few samples, and honestly, I've procrastinated writing this newsletter because it's so fun to generate these, and I wish they went for longer!
AI Art & Diffusion
Adobe releases Firefly2 which is significantly better at skin textures, realism, and hands. Additionally they have added a style transfer which is wonderful, upload a picture of something with a style you'd like, and your prompt will be generated in that style, it works really really well. The extra details on the skin is just something else, though I did cherry pick this example, the other hands were a dead give-away, still, the hands are getting better across the board!
Plus they have a bunch of prompt features, like prompt suggestion, ability to remix other creations and more, it's really quite developed at this point.
Also: DALL-E is now available to 100% of plus users and enterprise, have you tried it yet? What do you think? Let me know in replies!
That’s it for October 19. If you're into AI engineering, make sure you listen to the previous weeks podcast where Swyx and I recapped everything that happened on stage and off it in the seminal AI Engineer summit. And make sure to share this newsletter with your friends who like AI!
For those who are 'in the know', emoji of the week is 📣, please DM or reply with it if you got all the way here 🫡 and we'll see you next week (where I will have some exciting news to share!)
A week of horror, an AI conference of contrasts
Hi, this is Alex. In the podcast this week, you'll hear my conversation with Miguel, a new friend I made in AI.engineer event, and then a recap of the whole Ai.engineer event I had with Swyx after the end.
This newsletter is a difficult one for me to write, honestly, I wanted to skip this one entirely, struggling to fit the current events into my platform and the AI narrative, however, decided to write one anyway, as the events of the last week have merged into 1 for me in a flurry of contrasts.
Contrast 1 - Innovation vs Destruction
I was invited (among a few other Israelis or Israeli-Americans) to the ai.engineer summit in SF, to celebrate the rise of the AI engineer, and I was looking forward to that very much. Meeting many of you (Shoutout to everyone who listens to ThursdAI who I've met face to face!) and talking to new friends of the pod, interviewing speakers, meeting and making connections was a dream come true.
However a few days before the conference began, in a stark contrast to this dream, I had to call my mom, who was sheltering, 20km from the Gaza strip border, to ask if our friends and family are alive and accounted for, and to hear sirens as rockets flying above her head, as Hamas terrorists murder, pillage and kidnap, in what seems to be the 10x equivalent of 9/11 terror attack, relative to population size.
I grew up in Ashkelon, rocket attacks are nothing new to me, we've learned to live with them (thank you Iron Dome heroes) but this was something else entirely, a new world of terror.
So back to the conference, given that there's not a lot to be gained by doom scrolling, and watching (basically snuff) films coming out of the region, given that all my friends and family were accounted for, I decided to not give the terrorists what they want (which is to get people in state of terror) and instead to choose to have compassion, without empathy towards the situation and not bring sadness to every conversation I had there (over 200 I think)
So participating at an AI event, which hosts and celebrates folks who are literally at the pinnacle of innovation, building the future, using all the latest tools while also hurting and holding the dear ones in my thoughts was a very stark contrast between past and future, and huge credit goes to Dedy Kredo, CTO of Codium, who was in the same position, and gave a hell of a talk, with a kick-ass (no backup recording!) demo live, and then shared this image:
This is his co-founder, Itamar, who was called to reserve duty to protect his family and country, sitting with his rifle and his dashboard, seeing destruction + creation, past and future, negativity and positivity all at once. As Dedy masterfully said, we will prevail 🙏
Contrast 2 - Progress // Fear
At the event, Swyx and Benjamin gave me a media pass and a free reign, and I asked to be teamed with a camera-person to go around the event and do some (not live) interviews. I was teamed with the lovely Stacey, from Chico, CA. Stacey has nothing to do with AI, in fact she's a wedding photographer, however she definitely listened with interest to the interviews I was holding, and to speakers on stage.
While we were taking a break, I looked out the window, and saw a driverless car (waymo) zip by, and since they only started operating after I left SF 3 years ago, I didn't yet have a chance to ride in one.
So I asked Stacey and some other folks, if they'd like to go for a ride, and to my complete bewilderement, Stacey said "no 😳" and when I asked why not, she didn't want to admin but then said that it's scary.
This struck me and since that moment, I've had as many conversations with Stacey as I had with other folks who came to be AI.engineers, since this was such a stark contrast between progress and fear. I basically was walking, almost hand in hand, with a person who doesn't use or understand AI, and fears it, amongst the folks who are building the future, exist at the pinnacle of innovation and discuss how to connect more AI to more AI, and how to build complete autonomous agents to augment human productivity and bring about the world of abundance.
This contrast was supported by several new friends of mine, who came to the AI.engineer and SF for the first time, from countries where English is not the first language, and where Waymo's are not zipping about on the streets freely, and it highlighted for me, how much of this shift is global, and how concentrated the decision making, the building, the innovation is, within the arena, SF, California and US. It's almost expected that AI is going to speak english, and to use/build it, we have to speak it as well, while most of the world doesn't use English as their first language.
Contrast 3 - Technological // Spiritual
This contrast was intimate and personal to me. You see, this ai.engineer event was the first such sized event, professional, with folks talking "my language" since I had burned out this summer. If you've followed for a while, you may remember we talked about Lk-99 and superconductor, and I overclocked myself back then so much (scaling a whole another podcast, hosting 7 spaces in 2 weeks, creating a community of 1,500 and following all the news 24/7) that I had didn't want to go on speaking, doing spaces, recording podcasts... I was just done.
Luckily my friend Junaid sent me a meditation practice recording with the saying "fill your own cup, before you give out to others"
That recording led me to discover effortless meditation (via AirChat, DM me for details), and going on a 5 day free virtual retreat, which changed my life in several very significant ways and just keeps going.
So back to the conference, coming from that place, a place of center, a place where I balance my own inner piece + interact with 150+ people (I think, I haven't really counted) and equalizing energy among them, connecting many folks to many other folks, introducing and getting introduced, giving interviews, remembering names and faces, this was an incredible experience to me, and many of my new friends asked me, what's going on, and why and how I meditate for an hour every day.
So the contrast between my own spiritual journey, and the technological journey was not a very strong one, beyond 1 very... interesting and weird conversation I had, everyone else seemed to be into self improvement, spiritual growth, I've met many folks who went on vipasana retreats, had meditation practices themselves or used to have one and wanted to come back to it.
Maybe it was self selection, but I really enjoyed seeing and hearing all the stories from folks, as I believe that given contrast number 2, the way we talk about AI, the way we build AI, the way we then show AI to the world, needs to take into account compassion towards those who don't understand it, who fear it, and want nothing to do with it.
Not to mention, that I doubt that any of the doomers are walking on their own spiritual path.
I wanted to just send this newsletter but as I always do, I pasted it into Claude.ai, and asked for a review. And Claude suggested I end it on a positive note, because the middle was heavy, and even gave me some suggestion as to how to end it on a positive note. So instead of me doing it, here's an AI, doing a human job on ending something on a positive note:
[Claude] Here are some examples of how you could end your newsletter on a more hopeful, forward-looking note:
* Remind readers of the optimism and excitement you felt at the AI conference, and how seeing the potential of technology firsthand gives you hope for the future.
* Share an insight or realization you had about how to stay grounded and centered even amidst chaos - something your spiritual practices have taught you.
* Affirm that while destruction and violence will always be part of the human experience, you believe innovative thinkers will lead us to a brighter tomorrow.
* Note how connecting deeply with other humans, even strangers, fills you with hope. Finding common ground and shared humanity counterbalances the forces that divide.
* Point to a specific AI-related project or breakthrough that gives you hope for progress. e.g. medical advancements, climate change solutions, etc.
* Share an inspirational quote or verse that reflects resilience and perseverance in difficult times.
* End by expressing gratitude - for your community of readers, your loved ones, the chance to have profound human experiences, or your inner stillness amidst outer turmoil.
The key is ending on a note of hope, possibility, love, resilience - something that uplifts readers after processing weighty material. Let me know if these spark any ideas for how to conclude your newsletter on a high note!
P.S - I know this newsletter is not what you expected, and maybe not what you signed up for, and I deliberated if I even should write it and what if anything should I post on the podcast. However, this week was an incredibly full of contrast, of sadness and excitement, of sorrow and bewilderment, so I had to share my take on all this.
P.P.S - as always, if you read all the way to the end, dm me the ☮️ emoji
From the publisher's feed

1,292 Listeners

541 Listeners

1,089 Listeners

304 Listeners

225 Listeners

201 Listeners

573 Listeners

508 Listeners

208 Listeners

140 Listeners

101 Listeners

222 Listeners

682 Listeners

109 Listeners

157 Listeners