
Sign up to save your podcasts
Or


Companies competing in the chatbot wars are using something known in the industry as “the Pile” to train their large language models. It’s a trove of open-source data made up of text scraped from all around the internet, including Wikipedia and the European Parliament. Annie Gilbertson, investigative reporter for Proof News, recently took a deep dive into the Pile and discovered something else: a dataset called “YouTube Subtitles.” Marketplace’s Lily Jamali spoke with Gilbertson about her investigation and how YouTube creators feel about their content being used without their consent.
By Marketplace4.4
7777 ratings
Companies competing in the chatbot wars are using something known in the industry as “the Pile” to train their large language models. It’s a trove of open-source data made up of text scraped from all around the internet, including Wikipedia and the European Parliament. Annie Gilbertson, investigative reporter for Proof News, recently took a deep dive into the Pile and discovered something else: a dataset called “YouTube Subtitles.” Marketplace’s Lily Jamali spoke with Gilbertson about her investigation and how YouTube creators feel about their content being used without their consent.

30,636 Listeners

8,794 Listeners

936 Listeners

1,389 Listeners

1,288 Listeners

3,229 Listeners

1,719 Listeners

9,733 Listeners

1,648 Listeners

5,483 Listeners

113,272 Listeners

1,450 Listeners

9,548 Listeners

10 Listeners

35 Listeners

5,593 Listeners

16,489 Listeners