
Sign up to save your podcasts
Or


Companies competing in the chatbot wars are using something known in the industry as “the Pile” to train their large language models. It’s a trove of open-source data made up of text scraped from all around the internet, including Wikipedia and the European Parliament. Annie Gilbertson, investigative reporter for Proof News, recently took a deep dive into the Pile and discovered something else: a dataset called “YouTube Subtitles.” Marketplace’s Lily Jamali spoke with Gilbertson about her investigation and how YouTube creators feel about their content being used without their consent.
By Marketplace4.5
12561,256 ratings
Companies competing in the chatbot wars are using something known in the industry as “the Pile” to train their large language models. It’s a trove of open-source data made up of text scraped from all around the internet, including Wikipedia and the European Parliament. Annie Gilbertson, investigative reporter for Proof News, recently took a deep dive into the Pile and discovered something else: a dataset called “YouTube Subtitles.” Marketplace’s Lily Jamali spoke with Gilbertson about her investigation and how YouTube creators feel about their content being used without their consent.

32,100 Listeners

30,666 Listeners

8,776 Listeners

934 Listeners

1,388 Listeners

1,657 Listeners

2,175 Listeners

5,461 Listeners

111,948 Listeners

56,508 Listeners

9,532 Listeners

10,282 Listeners

3,614 Listeners

6,089 Listeners

6,564 Listeners

6,381 Listeners

163 Listeners

3,000 Listeners

154 Listeners

1,395 Listeners

90 Listeners