After thousands of episodes, we've noticed DeepSeek has some persistent quirks — overused words, recycled analogy templates, and habits no system prompt seems to cure. What if we took a hundred scripts, wrote human feedback on each, and fine-tuned a version of DeepSeek optimized solely for producing this podcast? This episode breaks down the practical steps: collecting feedback data, choosing between supervised fine-tuning and DPO, structuring training examples within DeepSeek's 128K context window, and using LoRA to avoid catastrophic forgetting. We also tackle the question of where character personalities should live — baked into the fine-tune or kept in the system prompt. It's a deep dive into whether purpose-specific fine-tuning is practical engineering or just a beautiful fantasy.
Episode #775107 — open it directly at myweirdprompts.com/775107