Elon Musk just called ChatGPT's training data "concerning." He's not wrong.
ChatGPT learned from 300 billion words scraped from the internet, including Reddit threads, Wikipedia articles, and news sites. But here's what most people miss: that training data cuts off in 2021, and OpenAI won't say exactly what's in it. Musk thinks this creates real problems around bias and misinformation that we're just starting to understand.
The numbers are pretty wild. Researchers found over 200 ways to bypass ChatGPT's safety filters, and OpenAI admits the system makes up information 15-20% of the time when asked factual questions. That's not a bug, it's how these models work. They predict the next most likely word, not necessarily the most accurate one.
In This Episode:
> Why ChatGPT's training data matters more than most people realize
> The specific examples Musk cited about political bias in AI responses
> What "hallucination" actually means and why it happens so often
> How prompt engineering can trick these systems into saying almost anything
Timestamps:
00:00 Introduction
01:30 What's actually in ChatGPT's training data
03:45 Musk's specific concerns about AI bias
06:20 The hallucination problem explained
08:15 Why safety filters don't really work
10:30 What this means for regular users
James breaks down the technical stuff without the Silicon Valley hype. If you're using ChatGPT for work or just curious about what's actually happening behind the scenes, this episode explains what Musk is really worried about.
🤖 AI moves fast. Follow Unboxed for daily episodes that keep you ahead of what's actually happening in artificial intelligence.
Learn more about your ad choices. Visit megaphone.fm/adchoices