Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Short answer: Reporting indicates that Anthropic used The Pile, a large text corpus containing the YouTube Subtitles dataset, for Claude. Apple research materials also described using The Pile, although Apple did not answer WIRED/Proof News questions about that investigation. This evidence concerns caption text—not proof that either company downloaded complete YouTube videos for that training run.

A separate 2026 lawsuit accuses Apple of obtaining videos through Panda-70M and bypassing YouTube protections. Those allegations, and Apple’s response invoking public availability, the DMCA and YouTube’s terms, have not been resolved by a merits ruling in the material available here.

What the 2024 reporting actually found

WIRED and Proof News reported in 2024 that the YouTube Subtitles dataset was included in The Pile. The Pile is a text corpus, so the relevant material was human-generated subtitle and caption text associated with YouTube videos.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The investigation identified 173,536 videos from more than 48,000 channels. That is the count reported by WIRED/Proof News, not a settled audit of every file or channel involved. A later amended complaint in Concord Music Group v. Anthropic gives a different figure—173,651 videos—showing why the numbers should be attributed rather than combined.

Nothing in that reporting by itself establishes that Apple or Anthropic downloaded the complete audiovisual versions of all those videos. Captions can be copied as text without establishing that the underlying video, audio track or channel metadata was used.

What Anthropic confirmed

Anthropic spokesperson Jennifer Martinez confirmed that Claude used The Pile and said, “The Pile includes a very small subset of YouTube subtitles.” That statement confirms the dataset connection while characterizing the YouTube material as a small part of the corpus.

What was reported about Apple

Apple research materials described use of The Pile, according to WIRED/Proof News. Apple did not respond to that investigation’s requests for comment. The reporting therefore links Apple’s published research references to The Pile, but it does not provide a company admission that Apple downloaded complete YouTube videos or explain precisely how any caption data was used in a particular model.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why creators were angry

Creators interviewed in the reporting said they had not been asked for permission or were frustrated to discover that their work appeared in the dataset. Their reactions describe consent and compensation concerns; they are not judicial findings that a company infringed copyright.

  • David Pakman, host of The David Pakman Show, said: “No one came to me and said, ‘We would like to use this.’”
  • Julie Walsh Smith, CEO of Complexly, said: “We are frustrated to learn that our thoughtfully produced educational content has been used in this way without our consent.”
  • Dave Farina, host of Professor Dave Explains, called for a conversation about compensation or regulation if AI products profit from creators’ work and threaten their livelihoods.

These comments explain the dispute’s human stakes: captions reflect years of scripting, teaching and production even when they are distributed as text. They do not, by themselves, establish what license applied to a particular caption or whether a court would find a legal violation.

The separate 2026 Apple lawsuit

Do not merge the caption-dataset story with the later Panda-70M case. In April 2026, Ted Entertainment and owners of MrShortGame Golf and Golfholics filed a proposed class action against Apple. The complaint alleges that their videos were accessed through Panda-70M and that their content appeared more than 500 times in that dataset.

What the complaint alleges

The plaintiffs describe Panda-70M as a video dataset and allege that Apple obtained or used YouTube material while circumventing YouTube protections. The complaint’s language is an allegation, not proof that the conduct occurred or that it was unlawful.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Apple’s response

As covered by MacRumors, Apple’s July 2026 response argued that the plaintiffs had made the videos publicly available and that access was permitted by the DMCA and YouTube’s terms. That is Apple’s litigation position. The sources reviewed do not establish a final ruling on the merits or a final disposition, so the lawsuit cannot accurately be described as proof that Apple broke the law.

How the two controversies differ

Question 2024 YouTube Subtitles / The Pile 2026 Apple / Panda-70M lawsuit
Material described Human-generated subtitle and caption text in a text corpus Video data and clips described in the plaintiffs’ complaint
Company connection Anthropic confirmed using The Pile; Apple research references described its use Apple is the defendant named in the proposed class action
Evidence status Investigative reporting, research references and Anthropic’s comment Plaintiffs’ allegations and Apple’s response; no merits ruling established here
Main issue Whether creator-produced caption text entered training data without creators’ awareness or consent Whether Apple unlawfully retrieved videos or bypassed YouTube protections
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What Apple’s current training-data policy says

In a September 9, 2026 disclosure, Apple described training data as a mixture of publicly available web-crawled information, directly licensed or purchased data, open-source data, user-study data and synthetic data. Apple said Applebot does not crawl sites requiring login credentials or protected by paywalls, and that it respects standard robots.txt directives that publishers can use to tell crawlers not to crawl or not to use website content for foundation-model training.

Apple also described filtering and processing steps. Those statements explain Apple’s general policy, but they do not establish the provenance or legality of The Pile, Panda-70M or any other specific third-party dataset in the disputes above.

What is established—and what is not

Established by the available record

  • The YouTube Subtitles dataset was included in The Pile.
  • Anthropic confirmed that Claude used The Pile and characterized the YouTube subtitles as a very small subset.
  • Apple research materials described use of The Pile, while Apple did not answer the reported investigation.
  • Creators reported learning that their material was included without being asked.
  • A separate 2026 complaint accuses Apple of conduct involving Panda-70M and source videos.

Not established by that record

  • That Apple or Anthropic each downloaded complete copies of every referenced YouTube video.
  • That every company associated with The Pile used it in the same way.
  • That inclusion of captions automatically violated copyright or a platform contract.
  • That the Panda-70M allegations have been proven in court.
  • That Apple’s general Applebot policy resolves the specific dataset disputes.

Why the wording matters

“Steal” is a characterization, not a legal finding supported by the material here. A precise account separates three questions: what files or text a dataset contained, which company used which dataset, and whether that use violated copyright, contract or anti-circumvention law. The 2024 reporting mainly addresses the first two questions for caption text. The 2026 lawsuit raises a different set of allegations about video retrieval and protections.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For viewers and creators, the practical issue is transparency: whether datasets identify their sources, provide a workable opt-out or licensing route, and explain how creator-produced material contributes to commercial AI systems. Those policy questions remain even when the legal outcome is unsettled.

The Bottom Line

The evidence supports a narrower conclusion than the headline: Anthropic confirmed using The Pile, which contained YouTube caption text, and Apple’s research referenced the same corpus. A later lawsuit separately alleges that Apple used Panda-70M to obtain videos improperly. Neither record, as presented here, proves that both companies copied complete YouTube videos or that either company has been found liable in court.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.