Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Short answer: MCP can connect an OpenAI model to external tools and services; the Responses API can accept PDFs for analysis; and the Images API can generate, edit, or vary images. Video is different: as of September 29, 2026, OpenAI’s Sora 2 models and Videos API are shut down, with no one-to-one replacement API available. MCP is not itself a video, PDF, or image generator—it gives a model access to tools.

What “ChatGPT MCP server” means in this workflow

MCP (Model Context Protocol) is a way for a model to use tools exposed by an MCP server. OpenAI’s developer guidance describes connecting models to remote MCP servers and local servers through Secure MCP Tunnel. A server can provide capabilities for connecting to and controlling an external service; the model may call those tools automatically or be required to get developer approval first.

That makes MCP a connection layer, not a media format or a standalone generation endpoint. The three jobs in this article are distinct:

  • Use an external tool: connect the model to an MCP server that exposes the relevant service or action.
  • Analyze a PDF: pass the document to the Responses API as file input.
  • Create or edit an image: use the Images API with a prompt, an input image, or both.

People often say “ChatGPT MCP server” to mean an MCP server connected to an OpenAI model. In a developer integration, distinguish the model or API client from the server: the server is the external tool provider, and the model decides whether a tool call is useful within the permissions and approval rules you configure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Connect a model to an MCP server

Choose the right connection type

For a publicly reachable service, configure the server with its server_url. For a private or on-premises service, OpenAI’s guidance describes using Secure MCP Tunnel with a tunnel_id. The provider may require OAuth authentication. The exact authentication and available tools depend on the MCP server you choose.

At a high level, the request configuration needs to identify the MCP server and establish whether tool calls need approval. The following is a conceptual checklist, not a complete API request: supply the public server’s server_url, or the private connection’s tunnel_id; configure authentication if the provider requires it; and choose an approval policy appropriate to the actions and data involved. Use the current OpenAI API reference and the server provider’s own instructions for the exact request fields and authentication flow before deploying.

Decide when the model may act

Automatic tool calls can make a workflow smoother, but they also let a model trigger actions on an external service without pausing for every call. An approval-gated policy gives a developer or user a chance to review sensitive actions first. Prefer explicit approval when a tool can publish, delete, purchase, change account settings, or transmit confidential material.

Before connecting, verify what the server’s tools can do and what information they receive. A tool that only reads public page metadata has a different risk profile from one that can send messages or modify records. Do not infer safety from the fact that a server speaks MCP.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Send a PDF to the Responses API

The Responses API accepts PDFs as input_file content items. The PDF can be supplied with a filename and the MIME type application/pdf, or by referencing a file ID. PDF processing can include both extracted text and page images, so a visually complex document may consume more tokens than its text alone suggests. A vision-capable model, such as GPT-4o or later, is needed for visual PDF parsing.

Request shape

The important content item has this structure:

{
  "type": "input_file",
  "filename": "report.pdf",
  "file_data": "data:application/pdf;base64,..."
}

This illustrates the PDF item, not a full authenticated request. Put it in the appropriate user message content for a Responses API request, alongside a clear instruction—for example, asking for the report’s findings, a table of specified values, or a comparison between named sections. If your file has already been uploaded and you have a file ID, use the file-ID form supported by the current API reference instead of embedding the data.

Respect size and visual-detail limits

  • A single PDF is limited to 50 MB.
  • The combined files in one request are also limited to 50 MB.
  • The PDF image detail setting can be auto, low, or high. Choose based on the task: small charts, fine print, or dense page layouts may need more visual detail than a request that only needs ordinary text.
  • Because processing may include page images as well as extracted text, account for additional token usage when sending long or image-heavy PDFs.

For a text-only extraction task, ask for specific fields and make the expected output format explicit. For a chart-reading or layout question, say which page or figure matters and request that the model distinguish visible evidence from inference. If accuracy matters, verify the cited page and values against the original document rather than treating a generated summary as a substitute for the PDF.

Generate, edit, or vary images with the Images API

The Images API supports three related operations: generating an image from a prompt, editing an input image, and creating variations. A prompt, an input image, or both can be used, depending on the operation. GPT image models return image data in base64 form; your client must decode that data and save it as an image file before a person or application can view it as a normal PNG, WebP, or JPEG.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose output controls deliberately

  • Format: documented options include png, webp, and jpeg. Choose based on whether you need broad compatibility, smaller web delivery, or a specific downstream workflow.
  • Quality: set the documented quality control to balance output fidelity and the requirements of your application.
  • Background: use the background control where the desired output calls for a particular background treatment.
  • Size: documented sizes include 1024x1024, 1024x1536, and 1536x1024. Pick the orientation that fits the intended placement instead of assuming every output should be square.

For an edit, describe both what must change and what should remain consistent. For a variation, provide the source image and state which qualities to vary. For generation, specify subject, composition, visual style, orientation, and any text or elements that must be included. These are prompt-writing practices, not guarantees that the model will reproduce every detail exactly.

The current OpenAI Images API reference does not specify an exact request schema or endpoint here, so no supposedly runnable OpenAI image command is reproduced. Use the current Images API reference for the operation-specific request fields, authentication, and base64 handling; those details should match the model and SDK version you deploy.

Can ChatGPT generate video now?

Not through the documented Sora 2 models and Videos API. OpenAI’s Videos reference states that Sora 2 models and the Videos API were shut down on September 24, 2026, and are no longer available; it also says there is no one-to-one replacement API. That is the status as of September 29, 2026. Older examples using /v1/videos are legacy documentation, not runnable instructions for a current integration.

This shutdown statement is specifically about the Sora 2 models and Videos API described in that reference. It does not establish that every video feature in every OpenAI product or third-party service has the same status. Check the service and product documentation for the exact capability you intend to use rather than assuming that an image-generation endpoint also generates video.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Security, privacy, and operational checks for MCP

OpenAI warns that remote MCP servers are third-party services not verified by OpenAI. A server may access, send, or receive data. Treat an MCP connection as a real integration with an outside provider, not as a harmless extension of the model.

  • Use a trusted provider-hosted server. Confirm who operates it, which tools it exposes, and what those tools do before connecting.
  • Require approval for sensitive actions. Avoid automatic execution where a call can disclose sensitive information or make consequential changes.
  • Review URLs returned by tools. A tool result can contain links or instructions that should not be followed blindly.
  • Log what you share. Keep an appropriate record of data sent to MCP servers, especially in workflows involving customer or internal information.
  • Defend against prompt injection. Treat text retrieved from external services as untrusted input. It can contain instructions intended to manipulate the model, so keep tool permissions narrow and do not let retrieved content override your application’s rules.
  • Check authentication and retention. OAuth may be required, and the server’s handling of submitted data depends on its own provider policies. Confirm those terms before sending documents or personal information.

Pick the appropriate path

Need Use Key consideration
Let a model interact with an external service MCP server connected to the model Connection type, OAuth if required, tool permissions, approval policy, and server trust
Ask questions about a PDF PDF as an input_file item in the Responses API 50 MB per file and combined request limit; page images can increase token use
Create or change a still image Images API generation, edit, or variation operation Choose format, quality, background, size, and provide an input image when needed
Generate video through the documented Sora 2 API Unavailable as of September 29, 2026 The Videos API shut down September 24, 2026; no one-to-one replacement API is available

Or skip the browser setup

If the task is to capture a web page as a screenshot—not to generate a new image, analyze a PDF, or create video—ScreenshotNeo is a separate website screenshot API and MCP server. Its one-request example captures a page as an image; see the ScreenshotNeo API documentation for options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify the page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Does MCP replace the Responses API or Images API?

No. MCP connects a model to external tools; PDF analysis and image generation are separate API workflows.

Can I use MCP to make a PDF or image without another service?

MCP itself does not generate media. A connected server can expose a tool that does so, but its capabilities depend on that server.

Does ScreenshotNeo generate images or videos?

No. ScreenshotNeo captures web pages as screenshots or PDFs; it is not a general image or video generation API.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.