Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Probably, in the broad sense that Sora appears to have encountered game-related video or imagery in training—but the public evidence does not establish which game footage was used, who supplied it, or whether a particular recording was copied. OpenAI describes broad categories of training data, not an itemized video list. Tests reported by TechCrunch and The Washington Post show Sora producing recognizable game-like scenes, but those outputs cannot identify a source file or prove that a publisher provided it.

What has OpenAI disclosed about Sora’s training data?

OpenAI’s 2025 Sora System Card says Sora was trained on a mixture of publicly available data, proprietary data accessed through partnerships, and custom datasets developed in-house. It describes public data from machine-learning datasets and web crawls, partnership data, and human feedback. OpenAI names Shutterstock and Pond5 as examples of partners, but does not identify them as sources of game footage.

That disclosure gives a broad picture of the data categories, not a video-by-video inventory. The Associated Press reported on February 15, 2024, that OpenAI had not disclosed the imagery and video sources used to train Sora. OpenAI’s general training explainer, published in 2026, also describes foundation models as using publicly available internet information, third-party partner data, and information supplied or generated by users, trainers, and researchers. That general explanation does not identify the contents or provenance of Sora’s specific training videos.

What evidence suggests Sora encountered game-related material?

Independent tests offer evidence of recognizable game-related behavior. TechCrunch reported that prompts including “Italian plumber game” produced game-like imagery and said game content may have found its way into Sora’s training data. The Washington Post reported in 2025 that Sora could generate clips resembling Minecraft, game logos, and a streamer playing Civilization. Researchers quoted by the Post said the results suggested versions of originals appeared in training data, while cautioning that resemblance alone does not show direct copying from a rights holder.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

These observations make exposure to game-related visual material a reasonable inference. They do not reveal whether a result was learned from a full gameplay recording, a short clip, a still image, a user-uploaded video, or material collected from another source. Nor do they establish that any particular company supplied it.

What can game-like output establish—and what can’t it?

Evidence or explanation What it supports What it does not establish
Recognizable gameplay, logos, or streamer scenes in generated clips Sora can produce game-related visual patterns; the behavior is consistent with exposure to related material. The identity of a source video, whether a complete recording was retained, or who supplied any training material. TechCrunch (2024) and The Washington Post (2025) reported output tests, not a chain of custody.
Publicly available internet material, partnership data, and custom datasets These are the broad data categories OpenAI says were used for Sora. OpenAI’s 2025 Sora System Card names Shutterstock and Pond5 as partnership examples. That a named partner supplied game footage, or that any particular upload was licensed or included.
A public upload containing game footage It could be one possible route by which game-related material became available online. That the upload was authorized by a rights holder, used in training, or directly supplied by a game publisher. No complete, auditable list of game videos in Sora’s training set is publicly identified in the cited accounts.

Generative output can reflect learned visual regularities without reproducing a source video verbatim. As Joanna Materzynska told The Washington Post, “The model is mimicking the training data. There’s no magic.” That observation explains why resemblance can be meaningful evidence of learned patterns without, by itself, identifying the precise material or route of access.

Does reproducing a game logo prove copyright infringement?

No. A recognizable logo or scene may raise questions about training and output, but it does not alone establish infringement. The relevant analysis depends on what was copied, how the material was obtained, what the output reproduces, and which jurisdiction’s copyright rules apply. It also matters whether a source was licensed, publicly accessible, or uploaded by a user; public availability does not itself prove permission.

TechCrunch quoted intellectual-property attorney Joshua Weigensberg saying, “Training a generative AI model generally involves copying the training data.” That is a general observation about model training, not a finding that a specific game video was used unlawfully or that a particular generated result infringes. The available reporting does not establish that Nintendo, Microsoft, Mojang, Twitch, or another named rights holder supplied game footage to OpenAI.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Can safeguards or provenance labels answer the training-data question?

Output filtering or other model safeguards may affect what Sora will generate, but a restriction on outputs is not a record of what was in its training data. Likewise, C2PA-style provenance metadata can help label or trace information about a generated media file; it does not establish which videos were used to train the model. Training-data provenance requires evidence about the data sources themselves, which the public disclosures described above do not provide at an itemized level.

What is the most accurate way to describe the evidence?

  • Supported: Sora can generate recognizable game-related visuals, and OpenAI says it trained Sora on broad categories that include public, partnership, and custom data.
  • Reasonable inference: The reported behavior is consistent with Sora having encountered game-related visual material.
  • Not established: Which game videos appeared in training, whether a particular clip was copied, whether a rights holder supplied it, or whether any use was unlawful.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.