Vanishing Gradients

Episode 26: Developing and Training LLMs From Scratch


Listen Later

Hugo speaks with Sebastian Raschka, a machine learning & AI researcher, programmer, and author. As Staff Research Engineer at Lightning AI, he focuses on the intersection of AI research, software development, and large language models (LLMs).
How do you build LLMs? How can you use them, both in prototype and production settings? What are the building blocks you need to know about?
​In this episode, we’ll tell you everything you need to know about LLMs, but were too afraid to ask: from covering the entire LLM lifecycle, what type of skills you need to work with them, what type of resources and hardware, prompt engineering vs fine-tuning vs RAG, how to build an LLM from scratch, and much more.
The idea here is not that you’ll need to use an LLM you’ve built from scratch, but that we’ll learn a lot about LLMs and how to use them in the process.
Near the end we also did some live coding to fine-tune GPT-2 in order to create a spam classifier!
LINKS
The livestream on YouTube (https://youtube.com/live/qL4JY6Y5pmA)
Sebastian's website (https://sebastianraschka.com/)
Machine Learning Q and AI: 30 Essential Questions and Answers on Machine Learning and AI by Sebastian (https://nostarch.com/machine-learning-q-and-ai)
Build a Large Language Model (From Scratch) by Sebastian (https://www.manning.com/books/build-a-large-language-model-from-scratch)
PyTorch Lightning (https://lightning.ai/docs/pytorch/stable/)
Lightning Fabric (https://lightning.ai/docs/fabric/stable/)
LitGPT (https://github.com/Lightning-AI/litgpt)
Sebastian's notebook for finetuning GPT-2 for spam classification! (https://github.com/rasbt/LLMs-from-scratch/blob/main/ch06/01_main-chapter-code/ch06.ipynb)
The end of fine-tuning: Jeremy Howard on the Latent Space Podcast (https://www.latent.space/p/fastai)
Our next livestream: How to Build Terrible AI Systems with Jason Liu (https://lu.ma/terrible-ai-systems?utm_source=vg)
Vanishing Gradients on Twitter (https://twitter.com/vanishingdata)
Hugo on Twitter (https://twitter.com/hugobowne)

This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit hugobowne.substack.com
...more
View all episodesView all episodes
Download on the App Store

Vanishing GradientsBy Hugo Bowne-Anderson

  • 5
  • 5
  • 5
  • 5
  • 5

5

12 ratings


More shows like Vanishing Gradients

View all
Odd Lots by Bloomberg

Odd Lots

2,001 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,468 Listeners

The a16z Show by Andreessen Horowitz

The a16z Show

1,101 Listeners

Talk Python To Me by Michael Kennedy

Talk Python To Me

581 Listeners

Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

Super Data Science: ML & AI Podcast with Jon Krohn

300 Listeners

Practical AI by Practical AI LLC

Practical AI

210 Listeners

Last Week in AI by Skynet Today

Last Week in AI

312 Listeners

Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

98 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

528 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

137 Listeners

Latent Space: The AI Engineer Podcast by Latent.Space

Latent Space: The AI Engineer Podcast

98 Listeners

The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

The AI Daily Brief: Artificial Intelligence News and Analysis

648 Listeners

Sharp Tech with Ben Thompson by Andrew Sharp and Ben Thompson

Sharp Tech with Ben Thompson

95 Listeners

High Signal: Data Science | Career | AI by Delphina

High Signal: Data Science | Career | AI

18 Listeners

OpenAI Podcast by OpenAI

OpenAI Podcast

61 Listeners