Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

PowerShell can convert a PDF to text by running a PDF extraction tool; it cannot do the conversion with Get-Content alone. One documented route is Apache PDFBox: use the export:text command with PDFBox 3.x, then read the resulting text file with PowerShell. PDFBox 2.x uses different command syntax. If the PDF is a scan made from images, ordinary text extraction may return little or no text; the cited PDFBox documentation does not establish an OCR workflow.

What you need before converting

Use a PDF parser such as Apache PDFBox 3.0 Command-Line Tools. PowerShell handles the surrounding work: it can launch Java, pass the PDF and output paths to PDFBox, check the process result, and read the text file afterward. Get-Content reads text files; it does not parse or convert PDFs.

  • Java must be installed and available as java in the shell’s path, or you must specify its full executable path.
  • Download the PDFBox application JAR for the release you intend to use, and substitute its real filename for the example name below.
  • Have a readable PDF and a destination folder where your account can create the text file.
  • Check the documentation or command help for the installed PDFBox release before relying on a particular option. PDFBox 3.x and 2.x do not use the same extraction command.

The commands below illustrate the documented PDFBox 3.x form. They are not a claim that a particular Java/PDFBox installation or PDF file has been tested.

Convert a PDF with PowerShell and PDFBox 3.x

Save this as a .ps1 file in the folder containing the PDFBox JAR and input.pdf, or update the paths to match your files. Replace pdfbox-app-3.y.z.jar with the actual JAR filename you downloaded. The 3.y.z notation is illustrative, not a filename to download.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
PowerShell for Sysadmins: Workflow Automation Made Easy
  • Book - powershell for sysadmins: workflow automation made easy
  • Language: english
  • Binding: paperback
  1. Set the paths and confirm that the required files exist.

    $jarPath = Join-Path $PSScriptRoot 'pdfbox-app-3.y.z.jar'
    $inputPath = Join-Path $PSScriptRoot 'input.pdf'
    $outputPath = Join-Path $PSScriptRoot 'output.txt'
    
    if (-not (Test-Path -LiteralPath $jarPath -PathType Leaf)) {
        throw "PDFBox JAR not found: $jarPath"
    }
    if (-not (Test-Path -LiteralPath $inputPath -PathType Leaf)) {
        throw "Input PDF not found: $inputPath"
    }
  2. Run the PDFBox 3.x text export command and stop if the Java process reports failure.

    & java -jar $jarPath export:text "-i=$inputPath" "-o=$outputPath"
    if ($LASTEXITCODE -ne 0) {
        throw "PDFBox failed with exit code $LASTEXITCODE"
    }

    The argument names and command form follow the PDFBox 3.x documentation: export:text, -i/--input, and -o/--output. Quoting the input and output arguments lets paths containing spaces be passed as single arguments. The call operator & invokes the executable in PowerShell.

  3. Check that an output file exists, then read it as one string.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
    if (-not (Test-Path -LiteralPath $outputPath -PathType Leaf)) {
        throw "PDFBox did not create the expected text file: $outputPath"
    }
    
    $text = Get-Content -LiteralPath $outputPath -Raw
    $text

    With -Raw, Get-Content returns the file contents as one string. Without it, PowerShell returns the contents as lines. To save the extracted text somewhere else, change $outputPath before running the export.

Use the command form for your PDFBox version

Do not mix the 3.x command with the older 2.x syntax. The official PDFBox 2.0 Command-Line Tools documentation shows ExtractText followed by options and the input filename, with an optional text-file argument. For example, the shape is:

java -jar .pdfbox-app-2.y.z.jar ExtractText [OPTIONS] <inputfile> [Text file]

Here, too, 2.y.z is a placeholder: use the JAR you actually have. Refer to that release’s help for its options and exact argument handling. A command written for PDFBox 3.x using export:text -i=... -o=... is not interchangeable with the 2.x form.

If you need to check which JAR you are invoking, inspect the filename and run the installed tool’s help rather than assuming that an older script matches your current download. Keep the command version and JAR version aligned.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose pages, encoding, and text order

The PDFBox 3.x command-line documentation describes options for page selection and sorting, as well as UTF-8 as the default output encoding. The precise option names and behavior can depend on the installed release, so consult its command help before adding options to a script. Do not copy an option from a different major version without checking it.

  • Extracting selected pages: Use the page-range controls documented for your release when you do not need the whole document. Verify the range syntax in that release’s help; do not assume a format from another PDF tool.
  • Text encoding: PDFBox 3.x documents UTF-8 as the default. If your downstream program expects another encoding, check the installed version’s documented encoding option and set it explicitly rather than guessing.
  • Reading order: Multi-column pages, sidebars, tables, and complex layouts can produce text in an order that differs from how the page looks. PDFBox documents sorting controls; try the appropriate release-specific option if reading order matters, and inspect the output against the page.
  • Markdown: PDFBox’s 3.0 command-line page notes Markdown output support since version 3.0.4. This is a version-qualified option, not a guarantee for every 3.x JAR; verify availability and syntax in the installed release’s help.
  • Password-protected PDFs: The 3.x documentation lists a password option. Use the installed release’s documentation for its exact syntax, and provide credentials only through a method appropriate to your environment. Avoid putting sensitive passwords in scripts that are shared or committed.

What if the PDF is scanned?

A scan may consist of page images rather than embedded text. A normal text extractor can only extract text represented in the PDF in a usable form; it should not be assumed to recognize words printed in an image. If the output is empty or missing the visible page content, first determine whether the PDF contains a text layer or only images.

The cited PDFBox command-line documentation establishes text extraction, but it does not establish an OCR method for image-only scans. Therefore this workflow alone cannot promise searchable or editable text from a scan. Use an OCR-capable workflow if recognition is required, and review its results for recognition errors before relying on names, numbers, or legal and financial text.

Run PDFBox from a reusable PowerShell function

If you will convert several files, put the invocation in a function that takes explicit paths. This keeps paths with spaces intact and checks the process exit code. The example uses PDFBox 3.x; update the JAR path and confirm the installed release’s syntax before use.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
function Convert-PdfToText {
    param(
        [Parameter(Mandatory)]
        [string] $PdfPath,
        [Parameter(Mandatory)]
        [string] $TextPath
    )

    $jarPath = Join-Path $PSScriptRoot 'pdfbox-app-3.y.z.jar'
    if (-not (Test-Path -LiteralPath $jarPath -PathType Leaf)) {
        throw "PDFBox JAR not found: $jarPath"
    }
    if (-not (Test-Path -LiteralPath $PdfPath -PathType Leaf)) {
        throw "Input PDF not found: $PdfPath"
    }

    & java -jar $jarPath export:text "-i=$PdfPath" "-o=$TextPath"
    if ($LASTEXITCODE -ne 0) {
        throw "PDFBox failed with exit code $LASTEXITCODE"
    }
    if (-not (Test-Path -LiteralPath $TextPath -PathType Leaf)) {
        throw "Output text file was not created: $TextPath"
    }
}

Convert-PdfToText -PdfPath '.Quarterly report.pdf' -TextPath '.Quarterly report.txt'

This function checks for missing files and a failed process. It does not validate whether the extracted text is complete or correctly ordered; open the output and compare it with the PDF when accuracy matters. For automation, decide how to handle existing output files and log failures in a way that fits your job rather than silently overwriting results.

Troubleshoot common failures

  • java is not recognized: Java is not available through the current process’s path, or it is not installed. Install/configure Java for your environment or invoke its executable by full path. Confirm the shell can run it before troubleshooting PDFBox.
  • JAR not found: The example filename is a placeholder, or the JAR is in another folder. Update $jarPath to the exact filename and location. Use Test-Path -LiteralPath $jarPath to check it.
  • Input or output path errors: Confirm the input exists and the output directory is writable. Keep paths quoted as in the examples, especially when they contain spaces. -LiteralPath makes PowerShell treat the file path literally when checking or reading a path.
  • Unknown command or option: Check whether the JAR is PDFBox 2.x or 3.x and use the matching command form. Consult that release’s help instead of combining options from different versions.
  • Output file is empty or nearly empty: The PDF may be image-only, encrypted, damaged, or structured in a way that does not yield useful text. Check whether the page has an extractable text layer; ordinary extraction is not evidence of OCR.
  • Text is scrambled or columns are interleaved: The PDF’s visual layout may not map cleanly to a linear text stream. Inspect the page-range and sorting options documented for your release, then check the resulting order manually.
  • PowerShell reports success but output is absent: Verify the output path and inspect the command’s process exit code. The examples check $LASTEXITCODE after invoking Java; avoid placing another native command between Java and that check.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and safe automation

Extraction time depends on the PDF and the machine; the available documentation does not establish a universal speed figure. For a single file, run the command interactively and inspect the text. For recurring jobs, capture errors, check the exit code, confirm an output was created, and test representative PDFs with the layouts and protections you expect to encounter.

Pass file paths as arguments rather than assembling a command string from uncontrolled input. Microsoft’s Start-Process documentation warns that untrusted data used with its FilePath parameter can create a security risk. If you use it to launch a process, keep the executable path under your control and handle arguments carefully. Direct invocation as shown above avoids needing Start-Process for this simple case.

Text extraction also changes the representation of a document: page geometry, fonts, and visual emphasis are not preserved in a plain text file. If later processing depends on document structure or visual placement, validate that the extracted result contains the information your next step needs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

ScreenshotNeo is a website screenshot API, not a PDF-to-text converter or OCR tool. It is relevant only if your actual input is a web page and you need a screenshot or PDF capture rather than extracted text. One GET request can return an image or PDF; see the ScreenshotNeo API documentation for parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie or consent banners and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status in headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000.

Sign up for ScreenshotNeo’s free plan to try website capture; it does not replace PDF text extraction.

For the PowerShell workflow, use the [Apache PDFBox command-line documentation](https://pdfbox.apache.org/3.0/commandline.html) as the reference for your installed release and verify the resulting text before using it downstream.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.