Practical guide
OCR PDF with a reviewable PDFWhirl workflow
OCR PDF processes one PDF containing scanned or image-based pages without modifying the file on your device. Output: One searchable PDF produced by OCRmyPDF with the selected recognition language.
What OCR PDF does
OCR PDF uses OCRmyPDF and Tesseract to add a searchable text layer to image pages while retaining a visual page representation.
OCR PDF runs as a queued server job after the source upload completes. The generated artifact receives a separate download link, so the original device file remains unchanged.
How to use this tool
Prepare the source
Choose the language that dominates the document and ensure scans are upright and readable. This preparation matters specifically before using OCR PDF.
Upload and inspect
Select one PDF containing scanned or image-based pages, wait for the OCR PDF upload to finish, and confirm that the displayed source is the intended file.
Choose the available options
Set only the OCR PDF controls needed for this output and review required fields before starting the job.
Download and verify
Search for several distinctive words and manually verify names, numbers, dates, amounts, and negations. Keep the source until the OCR PDF result has passed that review.
Common use cases
Prepare a delivery copy
Use OCR PDF to create a separate copy for a recipient without overwriting the maintained source.
Correct one document workflow
Apply OCR PDF to a known document problem after confirming the operation matches the intended outcome.
Create a review candidate
Generate a OCR PDF result that can be checked before it enters a records, publishing, or sharing process.
Supported inputs
- one PDF containing scanned or image-based pages
- One OCR PDF job at a time through the current public workspace
Output
One searchable PDF produced by OCRmyPDF with the selected recognition language.
Limitations to know
- Recognition is probabilistic; handwriting, low contrast, mixed languages, unusual fonts, and complex layouts can produce incorrect text.
- OCR PDF cannot reconstruct source information that is absent, damaged, or inaccessible in the uploaded document.
- Viewer, font, image, form, signature, and accessibility behavior can change after OCR PDF; compare important output with the source.
- No OCR PDF result should be treated as legal, archival, accessibility, or compliance certification without the required independent review.
Privacy and document handling
- OCR PDF uploads the selected source to configured object storage so the worker can process it.
- The worker attempts to remove temporary OCR PDF workspace files after the job, which is separate from stored input and output retention.
- Avoid using OCR PDF with confidential material unless the published storage and retention information is suitable for the document.
Stored input and output objects are deleted automatically: files uploaded without an account are removed about 24 hours after upload, and files belonging to a signed-in account are removed after 7 days. A scheduled sweep enforces this hourly against the object store itself. Workers attempt to remove per-job temporary directories after processing. This does not delete stored input or output objects.
Read the Privacy Policy and Security page for verified implementation details and open questions.
Troubleshooting
OCR PDF cannot start
Confirm the expected one PDF containing scanned or image-based pages finished uploading and every required OCR PDF option contains a valid value.
OCR PDF reports a processing error
Open the source locally, remove unsupported encryption where authorized, and retry OCR PDF with a smaller valid file.
OCR PDF output is not suitable
Search for several distinctive words and manually verify names, numbers, dates, amounts, and negations. If it still differs, retain the source and use a workflow designed for that document feature.
Still stuck? Review Help and common questions or contact support.
Frequently asked questions
Does OCR PDF replace my original file?
No. OCR PDF creates a separate stored output and does not edit the source on your device.
Is every OCR PDF result guaranteed to look identical?
No. OCR PDF depends on the source structure and processing engine, so important pages and features must be checked.
Should I delete my source after OCR PDF?
No. Keep the source until the OCR PDF output has been opened, compared, and accepted for its intended use.
How should I check my OCR PDF result?
Open every OCR PDF download and compare it with the unchanged source before sharing or deleting anything. Pay particular attention to this documented limitation: Recognition is probabilistic; handwriting, low contrast, mixed languages, unusual fonts, and complex layouts can produce incorrect text.