Skip to content

OCR PDF files

Make scanned PDFs searchable with text recognition.

or drop files here

Practical guide

OCR PDF with a reviewable PDFWhirl workflow

OCR PDF processes one PDF containing scanned or image-based pages without modifying the file on your device. Output: One searchable PDF produced by OCRmyPDF with the selected recognition language.

What OCR PDF does

OCR PDF uses OCRmyPDF and Tesseract to add a searchable text layer to image pages while retaining a visual page representation.

OCR PDF runs as a queued server job after the source upload completes. The generated artifact receives a separate download link, so the original device file remains unchanged.

How to use this tool

  1. Prepare the source

    Choose the language that dominates the document and ensure scans are upright and readable. This preparation matters specifically before using OCR PDF.

  2. Upload and inspect

    Select one PDF containing scanned or image-based pages, wait for the OCR PDF upload to finish, and confirm that the displayed source is the intended file.

  3. Choose the available options

    Set only the OCR PDF controls needed for this output and review required fields before starting the job.

  4. Download and verify

    Search for several distinctive words and manually verify names, numbers, dates, amounts, and negations. Keep the source until the OCR PDF result has passed that review.

Common use cases

Prepare a delivery copy

Use OCR PDF to create a separate copy for a recipient without overwriting the maintained source.

Correct one document workflow

Apply OCR PDF to a known document problem after confirming the operation matches the intended outcome.

Create a review candidate

Generate a OCR PDF result that can be checked before it enters a records, publishing, or sharing process.

Supported inputs

  • one PDF containing scanned or image-based pages
  • One OCR PDF job at a time through the current public workspace

Output

One searchable PDF produced by OCRmyPDF with the selected recognition language.

Limitations to know

  • Recognition is probabilistic; handwriting, low contrast, mixed languages, unusual fonts, and complex layouts can produce incorrect text.
  • OCR PDF cannot reconstruct source information that is absent, damaged, or inaccessible in the uploaded document.
  • Viewer, font, image, form, signature, and accessibility behavior can change after OCR PDF; compare important output with the source.
  • No OCR PDF result should be treated as legal, archival, accessibility, or compliance certification without the required independent review.

Privacy and document handling

  • OCR PDF uploads the selected source to configured object storage so the worker can process it.
  • The worker attempts to remove temporary OCR PDF workspace files after the job, which is separate from stored input and output retention.
  • Avoid using OCR PDF with confidential material unless the published storage and retention information is suitable for the document.

Troubleshooting

OCR PDF cannot start

Confirm the expected one PDF containing scanned or image-based pages finished uploading and every required OCR PDF option contains a valid value.

OCR PDF reports a processing error

Open the source locally, remove unsupported encryption where authorized, and retry OCR PDF with a smaller valid file.

OCR PDF output is not suitable

Search for several distinctive words and manually verify names, numbers, dates, amounts, and negations. If it still differs, retain the source and use a workflow designed for that document feature.

Still stuck? Review Help and common questions or contact support.

Frequently asked questions

Does OCR PDF replace my original file?

No. OCR PDF creates a separate stored output and does not edit the source on your device.

Is every OCR PDF result guaranteed to look identical?

No. OCR PDF depends on the source structure and processing engine, so important pages and features must be checked.

Should I delete my source after OCR PDF?

No. Keep the source until the OCR PDF output has been opened, compared, and accepted for its intended use.

How should I check my OCR PDF result?

Open every OCR PDF download and compare it with the unchanged source before sharing or deleting anything. Pay particular attention to this documented limitation: Recognition is probabilistic; handwriting, low contrast, mixed languages, unusual fonts, and complex layouts can produce incorrect text.

Learn more about this task

Browse all PDF guides

Related working tools

Browse all available tools