Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Short answer: To embed a real file in a PDF, add its bytes as an embedded-file stream and create a file specification that is indexed in the document catalog (or attached to a page). To extract those files, enumerate the PDF’s attachment mapping and write each payload to disk. Python’s pikepdf 10.15.0 documentation provides this interface. Do not confuse attachments with images, fonts, content streams, or XMP metadata: those are different PDF structures and require different extraction methods.

This guide answers “How do I embed a file in a PDF?”, “How do I extract attachments from a PDF?”, “How do I add arbitrary data to a PDF?”, and “How do I extract images from a PDF?” while showing the security, signing, archival, and forensic limits that matter in production.

What “embedded data” means in a PDF

A PDF is a graph of objects and streams. A stream can hold bytes, but not every stream is a user-facing attachment. A conventional attachment consists of two linked pieces:

  • Embedded-file stream: the actual payload bytes, such as a ZIP, JSON document, or binary file.
  • File specification: the PDF object that supplies a filename and points to the embedded stream.

For a document-wide attachment, the catalog’s Names dictionary can contain an EmbeddedFiles name tree mapping names to file specifications. This mechanism is defined for PDF 1.4 and later in the PDF Reference 1.7. A page-level file-attachment annotation uses a file specification at a particular page location and is commonly displayed as a paperclip icon.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

PDF 2.0 also supports Associated Files through the /AF relationship. It links a payload to a specific page, image, or other PDF object with a machine-readable relationship, rather than leaving the file as an undifferentiated document attachment. The PDF Association’s Application Note 002 describes this as a standardized, interoperable way to provide information related to a PDF object.

Structures that are not ordinary attachments

Structure What it stores Typical scope and extraction result
EmbeddedFiles name tree File specifications and embedded-file streams Document-wide downloadable files
File attachment annotation A file specification associated with a page location Page-level, often paperclip-visible
Associated File (/AF) A semantic relationship between a file and a PDF object Object-specific, machine-readable payload
XMP metadata Small descriptive properties in XML Document description, not a general file container; see Adobe’s XMP specifications
Image XObject, font, ICC profile, or content stream Resources needed to render a page Internal PDF data, not necessarily a user attachment

An image visible on a page is usually an Image XObject. PDF creation software may rescale or recompress it, so extracting that object can produce a valid image without reproducing the original source bytes. The PDF Association’s overview of files inside PDF also notes that rich-media, 3D, and other assets may use structures that a simple EmbeddedFiles listing does not enumerate.

How to embed arbitrary bytes with Python

Install the version of pikepdf you have approved for your application, then verify its API against the installed release. The examples below follow the documented pikepdf 10.15.0 interface; the documentation is evidence of the API, not a claim that this article executed the code.

Add an in-memory payload

import pikepdf

payload = b"binary data, JSON, or any other bytesn"

with pikepdf.Pdf.open("input.pdf") as pdf:
    pdf.attachments["payload.bin"] = payload
    pdf.save("output-with-attachment.pdf")

The assignment to pdf.attachments creates the attachment mapping and records the file specification in the catalog’s /AF array according to the pikepdf model documentation. Use a meaningful filename and choose an extension that helps downstream users identify the format. The bytes are opaque to PDF; pikepdf does not validate whether they are JSON, a ZIP archive, or executable content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Attach an existing file

import pikepdf
from pikepdf import AttachedFileSpec

with pikepdf.Pdf.open("input.pdf") as pdf:
    spec = AttachedFileSpec.from_filepath(pdf, "data/report.json")
    pdf.attachments["report.json"] = spec
    pdf.save("output-with-report.pdf")

Keep the source path and the attachment name under your control. If multiple files use the same name, decide whether to reject, rename, or replace them; relying on viewer-specific duplicate-name behavior is unsafe. For passwords or encrypted inputs, supply the appropriate opening credentials and confirm that your save preserves the required encryption policy.

Attach data with an explicit relationship

When a file belongs to a particular page, image, invoice line, or other object, use an Associated File relationship rather than treating it as a generic document download. The exact object construction depends on the PDF library and the relationship value required by your conformance profile. Validate the resulting PDF with a PDF 2.0 or PDF/A-aware validator; a generic attachment mapping alone does not prove that the required /AF relationship is present or semantically correct.

How to extract attachments from a PDF

For conventional attachments exposed by pikepdf, iterate Pdf.attachments and call read_bytes():

import os
import re
import pikepdf


def safe_name(name: str) -> str:
    # Prevent path traversal and platform-specific surprises.
    name = os.path.basename(name)
    name = re.sub(r"[^A-Za-z0-9._ -]", "_", name).strip(" .")
    return name or "attachment.bin"

with pikepdf.Pdf.open("input.pdf") as pdf:
    os.makedirs("extracted", exist_ok=True)
    for filename, attached_file in pdf.attachments.items():
        output_name = safe_name(str(filename))
        with open(os.path.join("extracted", output_name), "wb") as out:
            out.write(attached_file.read_bytes())

The documented interface is pdf.attachments['readme.txt'].read_bytes(); iterating the mapping generalizes that operation. In production, add limits for the number of entries and total uncompressed bytes, reject dangerous extensions when appropriate, and scan extracted files before opening them. Never trust an attachment filename as a filesystem path.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When the mapping is empty

  • The PDF may contain only page resources, not conventional attachments.
  • The file may use page annotations, Associated Files, rich-media, or 3D structures that your enumeration path does not expose.
  • The PDF may be malformed, encrypted, or damaged.
  • An incremental update may have marked an object deleted while older bytes remain physically present.

A normal viewer’s attachment panel is therefore not a forensic inventory. If the question is whether deleted or superseded payloads remain in earlier revisions, preserve the original bytes and use a revision-aware forensic workflow rather than rewriting the file first.

How to extract images from a PDF

Images are commonly Image XObjects, not attachments. A PDF-aware image extractor can identify those objects and decode their filters into image files. The result may be rescaled, color-converted, cropped, or recompressed compared with the source asset. If visual fidelity matters, render the page at a defined resolution and compare it with object extraction; if original-byte recovery matters, inspect the producer’s source files or revision history instead.

Do not expect XMP extraction to return a file. XMP is structured descriptive metadata, such as authoring or document properties. It can be reconciled with non-XMP PDF metadata, but it is not a replacement for an embedded-file stream.

Verification before distribution

  1. Open the saved PDF with a second independent reader and confirm that the attachment is listed and downloadable.
  2. Hash the original payload and the extracted payload with a cryptographic hash such as SHA-256; compare hashes when byte identity is required.
  3. Inspect the catalog and relevant page objects with a PDF object inspector to confirm whether the file is document-level, page-attached, or Associated.
  4. Validate PDF/A or PDF 2.0 conformance if an archive, accessibility, or exchange standard requires it.
  5. Test a signed workflow separately. Re-saving a signed PDF normally invalidates its signature, and removing or changing an attachment can alter the signed byte range.

Troubleshooting common failures

“The attachment is not visible”

Check both the document-level EmbeddedFiles name tree and page attachment annotations. A viewer may hide unsupported Associated Files or rich-media. Inspect the PDF objects rather than assuming the file was never embedded.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“The extracted file is corrupt”

Confirm that you wrote binary bytes with "wb", not text mode. Compare hashes, check the attachment’s declared filters and compression, and test extraction with another PDF implementation. Do not manually decompress a stream that the library has already decoded.

“Saving fails or the output is much larger”

Malformed cross-reference data, encryption permissions, or incremental-update history can affect rewriting. Work on a copy, open with the correct password, and use a current pikepdf release. Rewriting can discard unreachable historical objects and can change signatures.

“The image does not match the original”

That is expected when the PDF creator resampled or recompressed the image. Extracting an Image XObject recovers the PDF representation, not necessarily the source photograph or design file.

“Removing attachments broke a workflow”

Attachments can be integral to digital-signing processes. pikepdf documents separate remove_attachments and remove_external_access operations in its sanitization documentation. Treat removal as a policy decision, preserve the original, and revalidate signatures and archival conformance afterward.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your workflow first needs a clean screenshot of a web page before placing it into a PDF, ScreenshotNeo provides a single API call instead of maintaining browser automation. It removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, failed loads, and cache hits are not billed. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000.

cURL (see the ScreenshotNeo documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Create a free ScreenshotNeo account to use the 1,000-shot monthly allowance without a card.

Frequently Asked Questions

Can an embedded PDF file be encrypted separately from the PDF?

The PDF and its embedded stream are governed by the document’s encryption and permissions model; a separate payload-level encryption scheme requires encrypting the bytes before embedding and managing that key independently.

Does renaming an attachment change its contents?

No. Renaming changes the file specification’s name, while the embedded stream bytes remain unchanged; verify both the displayed name and payload hash in workflows that depend on identity.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can a PDF contain arbitrary binary bytes?

Yes, an embedded-file stream can carry arbitrary bytes, but consumers still need a format definition, safe handling policy, and suitable limits.