
Sign up to save your podcasts
Or


Pretraining and post-training chapter of Deep Dive into LLMs like ChatGPT by Andrej Karpathy.
In this section, the speaker explores the concept of base model inference, explaining how large AI models are trained, released, and function as token simulators rather than full assistants. The key points discussed include:
Base Models and Their Availability
* Training large AI models is extremely costly, but big tech companies often release “base models“ after training.
* A base model is a token simulator that predicts text sequences but is not yet an AI assistant.
Examples of Base Models
* GPT-2 (1.5B parameters, trained on 100B tokens) was one of the first widely released base models.
* LLaMA3 (405B parameters, trained on 15T tokens by Meta) is a modern, larger base model.
Components of a Model Release
* Requires two main parts:
* Python Code
* Model Parameters
Base Model Behavior
* It functions as an advanced autocomplete system, generating text based on statistical patterns from training data.
* It does not inherently provide factual or structured responses like an assistant.
Characteristics of Base Models
* Stochastic Nature - Given the same input, different competitions may be generated.
* Knowledge Compression - Acts like a lossy “zip file” of internet text, storing probabilistic patterns rather than explicit facts.
* Memorization & Regurgitation - Can recall high-frequency training data, sometimes verbatim (e.g, Wikipedia entries).
Limitations of Base Models
* Cannot provide factual updates beyond their training data cutoff.
* Tends to “hallucinate” (generate plausible but false information)
Practical Uses of Base Models
* In-context Learning - Few-shot prompting enables them to recognise and follow simple patterns
* Simulated AI assistants - Carefully structured prompts can trick a base model into behaving like an assistant by mimicking a conversation format
Acknowledgment
The Neural Network Internals and Training of Angrej Karpathy's video
The video clip of Deep Dive into LLMs Like ChatGPT by Andrej Karpathy
This is chapter3 from Andrej Karpathy video
Understanding Reasoning Model
The original article is https://www.linkedin.com/pulse/understanding-reasoning-large-language-models-bowen-li-kgadc/?trackingId=WtLIWBUiTFmVwCPL4HCwUw%3D%3D
The audio generated by Amazon Polly. It is not better than the OpenAI ChatGPT audio service.
From the publisher's feed