iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Short answer: Reporting indicates that Anthropic used The Pile, a large text corpus containing the YouTube Subtitles dataset, for Claude. Apple research materials also described using The Pile, although Apple did not answer WIRED/Proof News questions about that investigation. This evidence concerns caption text—not proof that either company downloaded complete YouTube videos for that training run.
A separate 2026 lawsuit accuses Apple of obtaining videos through Panda-70M and bypassing YouTube protections. Those allegations, and Apple’s response invoking public availability, the DMCA and YouTube’s terms, have not been resolved by a merits ruling in the material available here.
What the 2024 reporting actually found
WIRED and Proof News reported in 2024 that the YouTube Subtitles dataset was included in The Pile. The Pile is a text corpus, so the relevant material was human-generated subtitle and caption text associated with YouTube videos.
Free tools Windows power users keep installed
One-click scans. No signup required.
The investigation identified 173,536 videos from more than 48,000 channels. That is the count reported by WIRED/Proof News, not a settled audit of every file or channel involved. A later amended complaint in Concord Music Group v. Anthropic gives a different figure—173,651 videos—showing why the numbers should be attributed rather than combined.
#1 Best Overall
Nothing in that reporting by itself establishes that Apple or Anthropic downloaded the complete audiovisual versions of all those videos. Captions can be copied as text without establishing that the underlying video, audio track or channel metadata was used.
What Anthropic confirmed
Anthropic spokesperson Jennifer Martinez confirmed that Claude used The Pile and said, “The Pile includes a very small subset of YouTube subtitles.” That statement confirms the dataset connection while characterizing the YouTube material as a small part of the corpus.
What was reported about Apple
Apple research materials described use of The Pile, according to WIRED/Proof News. Apple did not respond to that investigation’s requests for comment. The reporting therefore links Apple’s published research references to The Pile, but it does not provide a company admission that Apple downloaded complete YouTube videos or explain precisely how any caption data was used in a particular model.
Why creators were angry
Creators interviewed in the reporting said they had not been asked for permission or were frustrated to discover that their work appeared in the dataset. Their reactions describe consent and compensation concerns; they are not judicial findings that a company infringed copyright.
Rank #3
- David Pakman, host of The David Pakman Show, said: “No one came to me and said, ‘We would like to use this.’”
- Julie Walsh Smith, CEO of Complexly, said: “We are frustrated to learn that our thoughtfully produced educational content has been used in this way without our consent.”
- Dave Farina, host of Professor Dave Explains, called for a conversation about compensation or regulation if AI products profit from creators’ work and threaten their livelihoods.
These comments explain the dispute’s human stakes: captions reflect years of scripting, teaching and production even when they are distributed as text. They do not, by themselves, establish what license applied to a particular caption or whether a court would find a legal violation.
The separate 2026 Apple lawsuit
Do not merge the caption-dataset story with the later Panda-70M case. In April 2026, Ted Entertainment and owners of MrShortGame Golf and Golfholics filed a proposed class action against Apple. The complaint alleges that their videos were accessed through Panda-70M and that their content appeared more than 500 times in that dataset.
Rank #4
What the complaint alleges
The plaintiffs describe Panda-70M as a video dataset and allege that Apple obtained or used YouTube material while circumventing YouTube protections. The complaint’s language is an allegation, not proof that the conduct occurred or that it was unlawful.
Apple’s response
As covered by MacRumors, Apple’s July 2026 response argued that the plaintiffs had made the videos publicly available and that access was permitted by the DMCA and YouTube’s terms. That is Apple’s litigation position. The sources reviewed do not establish a final ruling on the merits or a final disposition, so the lawsuit cannot accurately be described as proof that Apple broke the law.
How the two controversies differ
| Question | 2024 YouTube Subtitles / The Pile | 2026 Apple / Panda-70M lawsuit |
|---|---|---|
| Material described | Human-generated subtitle and caption text in a text corpus | Video data and clips described in the plaintiffs’ complaint |
| Company connection | Anthropic confirmed using The Pile; Apple research references described its use | Apple is the defendant named in the proposed class action |
| Evidence status | Investigative reporting, research references and Anthropic’s comment | Plaintiffs’ allegations and Apple’s response; no merits ruling established here |
| Main issue | Whether creator-produced caption text entered training data without creators’ awareness or consent | Whether Apple unlawfully retrieved videos or bypassed YouTube protections |
What Apple’s current training-data policy says
In a September 9, 2026 disclosure, Apple described training data as a mixture of publicly available web-crawled information, directly licensed or purchased data, open-source data, user-study data and synthetic data. Apple said Applebot does not crawl sites requiring login credentials or protected by paywalls, and that it respects standard robots.txt directives that publishers can use to tell crawlers not to crawl or not to use website content for foundation-model training.
Apple also described filtering and processing steps. Those statements explain Apple’s general policy, but they do not establish the provenance or legality of The Pile, Panda-70M or any other specific third-party dataset in the disputes above.
What is established—and what is not
Established by the available record
- The YouTube Subtitles dataset was included in The Pile.
- Anthropic confirmed that Claude used The Pile and characterized the YouTube subtitles as a very small subset.
- Apple research materials described use of The Pile, while Apple did not answer the reported investigation.
- Creators reported learning that their material was included without being asked.
- A separate 2026 complaint accuses Apple of conduct involving Panda-70M and source videos.
Not established by that record
- That Apple or Anthropic each downloaded complete copies of every referenced YouTube video.
- That every company associated with The Pile used it in the same way.
- That inclusion of captions automatically violated copyright or a platform contract.
- That the Panda-70M allegations have been proven in court.
- That Apple’s general Applebot policy resolves the specific dataset disputes.
Why the wording matters
“Steal” is a characterization, not a legal finding supported by the material here. A precise account separates three questions: what files or text a dataset contained, which company used which dataset, and whether that use violated copyright, contract or anti-circumvention law. The 2024 reporting mainly addresses the first two questions for caption text. The 2026 lawsuit raises a different set of allegations about video retrieval and protections.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchFor viewers and creators, the practical issue is transparency: whether datasets identify their sources, provide a workable opt-out or licensing route, and explain how creator-produced material contributes to commercial AI systems. Those policy questions remain even when the legal outcome is unsettled.
The Bottom Line
The evidence supports a narrower conclusion than the headline: Anthropic confirmed using The Pile, which contained YouTube caption text, and Apple’s research referenced the same corpus. A later lawsuit separately alleges that Apple used Panda-70M to obtain videos improperly. Neither record, as presented here, proves that both companies copied complete YouTube videos or that either company has been found liable in court.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

