Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Former OpenAI researcher Suchir Balaji argued that training generative AI on copyrighted works raises serious fair-use questions, especially when commercial products may compete with the works they learn from. His argument is not a court ruling: Balaji’s own essay says fair use must be decided case by case, and the sources behind the October 2024 report do not settle whether OpenAI’s practices are lawful.

What did Suchir Balaji allege about his work at OpenAI?

In an October 24, 2024 report, Futurism, citing a New York Times interview, described Balaji as a former OpenAI researcher who worked at the company for four years. The report said he was among the staff who collected and organized web-gathered data used to train large language models.

According to the report, Balaji’s view changed as ChatGPT became a commercial product following its November 2022 release. He became concerned that AI products could produce material reflecting or mimicking copyrighted source works. Futurism quoted him saying, “If you believe what I believe, you have to just leave the company.” He also called the model “not a sustainable model” for “the internet ecosystem as a whole.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What was Balaji’s fair-use argument?

In his October 23, 2024 essay, “When does generative AI qualify for fair use?”, Balaji starts from the premise that training generative models involves copying copyrighted data. He examines that premise through the four statutory fair-use factors: the purpose and character of the use, including whether it is commercial; the nature of the copyrighted work; the amount and substantiality used; and the effect on the potential market for or value of the work.

Balaji argues that a commercial AI system’s outputs may substitute for original works, and that the existence of a market for data licenses is relevant to potential market harm. These are his arguments about how the factors apply—not findings by a court. He expressly cautions: “Because fair use is determined on a case-by-case basis, no broad statement can be made about when generative AI qualifies for fair use.”

Why the market question is difficult to answer

Balaji acknowledges that the training data is not publicly known and that effects on markets vary by source. As a result, his essay says the market-impact question cannot be answered directly for every work. The argument identifies a concern to assess, but does not establish that every work used in training is harmed or that every AI output substitutes for an original.

How did OpenAI respond?

Futurism reported that OpenAI told The New York Times it builds its “AI models using publicly available data, in a manner protected by fair use and related principles” and called this “critical for ‘US competitiveness.’” That is the company’s reported position. The account does not resolve whether the use of any particular work qualifies as fair use.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What does the disagreement turn on?

Balaji’s essay and OpenAI’s reported response leave several questions for a case-specific legal analysis. The following are the points of disagreement to examine, not conclusions about how a court will rule:

  • Authorization: Whether copying a work for training is authorized or otherwise permitted.
  • Purpose and character: Whether the training use is sufficiently distinct or transformative, and how its commercial nature bears on the analysis.
  • Amount and outputs: How much of a work is copied and whether a system’s outputs reflect or reproduce protected material.
  • Market effect: Whether AI products substitute for original works or affect the market for those works and for licenses to use them.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the report and essay establish—and what they do not

Together, the October 2024 report and Balaji’s essay establish that he objected to OpenAI’s approach to training data and argued that commercial generative AI can raise substantial fair-use concerns. They also establish that OpenAI asserted a fair-use basis for using publicly available data. They do not establish that a court has found OpenAI liable for infringement, or that training on copyrighted material is categorically fair use.

The Futurism article refers to copyright lawsuits but does not provide a current disposition of those cases. It should therefore be read as a report of Balaji’s allegations and the competing positions, not as an update on the present status of litigation.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.