Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Yes, Codex can participate in workflows that produce videos, PDFs, and images through MCP, but MCP is the connection layer rather than a universal renderer. Run codex mcp-server when an MCP-compatible client needs Codex as a callable stdio tool. Use Codex document skills for formatted PDFs, the Responses API image-generation tool for raster images, and the asynchronous Videos API for MP4 output. The Codex App Server is a better integration when you need thread lifecycle events, streaming progress, and diff updates.

What the Codex MCP server actually does

OpenAI documents the operational model plainly: run codex mcp-server and connect from any MCP client that supports stdio servers (OpenAI engineering documentation). The MCP server exposes Codex capabilities to the client; it does not turn MCP itself into a PDF renderer, image model, or video encoder.

Your workflow normally has three layers:

  • Client and transport: Claude, Cursor, another MCP client, or a custom host starts Codex over local stdio.
  • Codex session: Codex interprets the request, gathers inputs, calls tools, and writes or returns files according to the permissions you grant.
  • Asset producer: document skills create or revise PDFs, the Responses API image tool generates or edits pixels, and the Videos API runs an asynchronous video job.

That separation matters for reliability. If a request says “make a PDF,” Codex still needs the source content, layout constraints, a writable location, and a review step. If it says “make a video,” the API returns a job that must be checked and downloaded after rendering.

Choose MCP or the Codex App Server

Decision point codex mcp-server Codex App Server
Primary purpose Make Codex a callable tool for an existing MCP workflow. First-class integration for richer Codex sessions.
Transport Local stdio. Application integration with session semantics.
Session features Only the operations exposed through MCP endpoints. Thread lifecycle, streaming progress, and diff updates.
Best fit A client that already starts MCP servers and needs Codex in its tool graph. An application that owns conversation state, progress UI, and code or file diffs.

Use MCP when portability and a small callable surface are more important than session control. Choose the App Server when your host must display incremental progress, manage threads, or apply diffs as first-class events.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Set up a local MCP connection

  1. Install and authenticate Codex. Verify that the codex command runs in the account and environment where your MCP client will launch it.
  2. Start the server from the client configuration. Configure the client to execute codex mcp-server as a stdio server. The exact JSON or UI field names differ by client, so keep the command and working directory explicit.
  3. Test with a harmless task. Ask Codex to create a text file in a temporary directory and return the path. Confirm that the client shows the tool call and that your approval policy behaves as expected.
  4. Add documentation access only when needed. OpenAI provides a read-only documentation MCP endpoint at https://developers.openai.com/mcp. The Codex CLI command is:
codex mcp add openaiDeveloperDocs --url https://developers.openai.com/mcp

Keep the documentation server read-only. It can supply current reference material, while the local Codex server remains responsible for the session and approved file operations.

Generate a formatted PDF with Codex

The Codex app includes document skills for reading, creating, and editing PDF files with professional formatting and layouts (OpenAI’s Codex app announcement). The practical flow is:

  1. Provide the source material: Markdown, plain text, a spreadsheet, or an existing PDF to revise.
  2. State the audience, page size, margins, typography, heading hierarchy, tables, images, and accessibility requirements.
  3. Ask Codex to create or revise the formatted PDF and save it to a named output path.
  4. Have Codex inspect the generated file, then open it yourself to check page breaks, clipped content, fonts, links, image resolution, and the final page count.

A useful request is: “Create a US-letter PDF from brief.md. Use 1-inch margins, a title page, numbered headings, a table of contents, and page numbers. Keep tables together where possible, save the result as build/report.pdf, and inspect the exported file for overflow before returning it.”

Do not assume a particular PDF library or visual result. The skill and the tools available in your Codex environment determine the implementation. Treat the first export as a draft and review the actual file.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Generate or edit images through the Responses API

The Responses API has a native image-generation tool that can create a new image or edit an existing one. It supports PNG, WebP, and JPEG output, quality values of low, medium, high, or auto, sizes of 1024x1024, 1024x1536, and 1536x1024, and optional partial-image streaming (Responses API reference). OpenAI lists GPT Image 1 and GPT Image 1 mini in its model catalog (model catalog).

Python request

Set OPENAI_RESPONSES_MODEL to a Responses model enabled for image generation. The script sends a prompt, requests a 1536×1024 high-quality image, and writes the base64 result when the response contains an image-generation item.

import base64
import json
import os
import requests

model = os.environ["OPENAI_RESPONSES_MODEL"]
payload = {
    "model": model,
    "input": "Create a clean editorial illustration of a developer orchestrating PDF, image, and video jobs through an MCP client.",
    "tools": [{"type": "image_generation"}],
    "tool_choice": "auto",
    "stream": False
}
headers = {"Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}", "Content-Type": "application/json"}
r = requests.post("https://api.openai.com/v1/responses", headers=headers, json=payload, timeout=180)
r.raise_for_status()
data = r.json()
for item in data.get("output", []):
    if item.get("type") == "image_generation_call" and item.get("result"):
        with open("illustration.png", "wb") as f:
            f.write(base64.b64decode(item["result"]))
        print("Wrote illustration.png")
        break
else:
    print(json.dumps(data, indent=2))

The response shape can evolve, so keep the raw JSON for diagnostics. For edits, provide the existing image through the input mechanism supported by your selected model and describe the exact changes; use a new output filename until the revision is approved.

Equivalent cURL request

curl https://api.openai.com/v1/responses 
  -H "Authorization: Bearer $OPENAI_API_KEY" 
  -H "Content-Type: application/json" 
  -d '{
    "model": "'"$OPENAI_RESPONSES_MODEL"'",
    "input": "Create a 1536x1024 editorial illustration of an MCP asset pipeline.",
    "tools": [{"type": "image_generation"}],
    "tool_choice": "auto"
  }'

For interactive agents, partial-image streaming lets the client show previews while the generation proceeds. Multi-turn edits are useful when the first composition is close but the subject, crop, or style needs correction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Generate a video as an asynchronous job

The Videos API accepts a text prompt and an optional input-reference image, then returns an asynchronous job. The documented models are sora-2 and sora-2-pro; clip lengths are 4, 8, or 12 seconds; supported sizes are 720x1280, 1280x720, 1024x1792, and 1792x1024. When rendering is complete, download the content, normally as an MP4 (Videos API reference).

Submit and poll with cURL

JOB=$(curl https://api.openai.com/v1/videos 
  -H "Authorization: Bearer $OPENAI_API_KEY" 
  -F model=sora-2 
  -F prompt="A slow dolly shot across a desk showing an MCP workflow producing a PDF, image, and video" 
  -F seconds=8 
  -F size=1280x720 | tee job.json | python -c "import sys,json; print(json.load(sys.stdin)['id'])")

while true; do
  STATUS=$(curl -s https://api.openai.com/v1/videos/$JOB 
    -H "Authorization: Bearer $OPENAI_API_KEY" | tee status.json | python -c "import sys,json; print(json.load(sys.stdin).get('status',''))")
  case "$STATUS" in
    completed) curl -L https://api.openai.com/v1/videos/$JOB/content -H "Authorization: Bearer $OPENAI_API_KEY" -o output.mp4; break;;
    failed) cat status.json; exit 1;;
  esac
  sleep 10
done

Python implementation

import os, time, requests

headers = {"Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}"}
with open("reference.png", "rb") as image:
    response = requests.post(
        "https://api.openai.com/v1/videos",
        headers=headers,
        files={"input_reference": image},
        data={
            "model": "sora-2",
            "prompt": "A product demo transition from an MCP tool call to a finished video asset",
            "seconds": "8",
            "size": "1280x720"
        },
        timeout=60
    )
response.raise_for_status()
job_id = response.json()["id"]
while True:
    status = requests.get(f"https://api.openai.com/v1/videos/{job_id}", headers=headers, timeout=30).json()
    if status.get("status") == "completed":
        video = requests.get(f"https://api.openai.com/v1/videos/{job_id}/content", headers=headers, timeout=180)
        video.raise_for_status()
        open("output.mp4", "wb").write(video.content)
        break
    if status.get("status") == "failed":
        raise RuntimeError(status)
    time.sleep(10)

Node.js implementation

const fs = require('fs');
const FormData = require('form-data');

const headers = { Authorization: `Bearer ${process.env.OPENAI_API_KEY}` };
const form = new FormData();
form.append('model', 'sora-2');
form.append('prompt', 'A product demo transition from an MCP tool call to a finished video asset');
form.append('seconds', '8');
form.append('size', '1280x720');
form.append('input_reference', fs.createReadStream('reference.png'));

const created = await fetch('https://api.openai.com/v1/videos', { method: 'POST', headers: { ...headers, ...form.getHeaders() }, body: form });
if (!created.ok) throw new Error(await created.text());
const job = await created.json();
while (true) {
  const status = await fetch(`https://api.openai.com/v1/videos/${job.id}`, { headers });
  const data = await status.json();
  if (data.status === 'completed') {
    const file = await fetch(`https://api.openai.com/v1/videos/${job.id}/content`, { headers });
    fs.writeFileSync('output.mp4', Buffer.from(await file.arrayBuffer()));
    break;
  }
  if (data.status === 'failed') throw new Error(JSON.stringify(data));
  await new Promise(resolve => setTimeout(resolve, 10000));
}

Use a durable job record containing the returned ID, prompt, model, duration, size, reference filename, and current status. Poll with backoff rather than assuming a fixed completion time. Download only after the status is completed, and retain the error payload for failed jobs.

Expose the workflow as your own MCP server

If you want another agent to request “make a PDF,” “generate an image,” or “render a video,” wrap each goal in a narrowly scoped MCP tool. OpenAI’s plugin guidance recommends recognizable user goals, schema validation, and explicit tool boundaries. The official TypeScript SDK is @modelcontextprotocol/sdk; the Python SDK is mcp. Streamable HTTP is available for networked servers (MCP server guidance).

Design the tools around outcomes

  • create_pdf: source content, layout constraints, output path, and a review flag.
  • generate_image: prompt, optional reference, output format, quality, and size.
  • start_video_job: prompt, optional reference, model, seconds, and size.
  • get_video_job: job ID and current status.
  • download_video: completed job ID and a controlled destination.

Return structured status, paths, and errors instead of burying them in prose. Keep network access and filesystem destinations constrained, and require approval before publishing or overwriting an existing asset.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Safety, permissions, and review

Codex is sandboxed by default and provides approval modes, network controls, and OpenTelemetry logging. Logged events can include prompts, tool-approval decisions, tool results, MCP-server usage, and network allow-or-deny decisions (OpenAI safety guidance).

  • Least privilege: give the session access only to the input and output directories it needs.
  • Approval gates: require confirmation before network calls, destructive file changes, or publication.
  • Reproducibility: save prompts, model choices, dimensions, durations, source references, and returned job IDs with the asset.
  • Human inspection: open every final PDF, inspect image dimensions and format, and watch a representative portion of each video.
  • Privacy: remove secrets and personal data from prompts, references, logs, and generated filenames.

For design handoff, OpenAI describes the Figma MCP Server as connecting Codex directly to Figma and tools such as Figma Make and FigJam (Figma partnership announcement). That can move a generated interface into an editable design review, but it does not replace checking the exported assets.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

The client cannot start Codex

Check that codex is on the executable path of the process launching the MCP client, not only in your interactive shell. Use an absolute working directory, verify authentication, and run codex mcp-server directly once to expose local configuration errors.

The MCP call hangs

Confirm that the client expects stdio rather than an HTTP URL, and that no wrapper is writing logs to standard output. Keep diagnostic logging on standard error so the protocol stream remains machine-readable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The PDF is clipped or visually inconsistent

Reduce layout complexity, specify page size and margins explicitly, and ask Codex to inspect the exported file. Check embedded fonts, long table rows, images that exceed the content box, and links that cross page breaks.

An image request returns text but no file

Inspect the raw Responses JSON for an image-generation output item, verify that the selected model supports the image tool, and ensure your code decodes the returned base64 data before writing the file. Keep the response JSON for diagnosis.

A video job fails or never completes

Validate the model, seconds, and size against the documented choices. Poll the job endpoint, handle a failed status, and retry only with a new job after recording the original error. Do not treat submission as proof that an MP4 exists.

Performance and operating-cost decisions

  • Use MCP for orchestration and tool portability; use the App Server when progress events and thread state are part of your product UI.
  • For images, choose the smallest acceptable size and quality, and stream previews when a human needs to iterate.
  • For video, select duration and resolution deliberately because every submitted job is asynchronous and produces a separate asset.
  • Cache source inputs and retain job metadata so a failed download does not force unnecessary regeneration.
  • Separate draft and production destinations, with a review gate before publication.

Or skip the browser setup

If your task is simply to capture a clean website image rather than generate new pixels, ScreenshotNeo provides a single-call website screenshot API and MCP server. It accepts consent banners like a visitor, removes more than 60 known consent platforms plus newsletter popups and chat widgets before capture, and lets you disable each cleanup step. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with the result identified by X-Page-Verdict and X-Billed headers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cURL (see the ScreenshotNeo API docs):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Frequently Asked Questions

Can one MCP tool call stream an image preview and a video render at the same time?

They are different lifecycles. Image generation can optionally stream partial previews through the Responses API, while a video request remains an asynchronous job that you poll and download after completion.

Is a custom MCP server required to use Codex with an existing MCP client?

No. Start the documented local codex mcp-server stdio server. Build a separate MCP server only when you need to expose your own narrowly scoped PDF, image, or video tools.

Where should generated files be reviewed before publication?

Use a controlled output directory, preserve the prompt and job metadata, and require a human inspection step for the actual PDF, image, or video before moving it to a published location.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.