H
HidePDF
Redact PDFs in your browser. Nothing leaves your device.
100% local · fully local

Does Redacting an OCR'd PDF Remove the Hidden Text Layer?

A scanned PDF that was OCR'd looks like a picture — but copy-paste still works because a text layer sits beneath the image. Redaction that only paints over the scan leaves that layer intact. You must remove or destroy the text you cannot see.

PDF
Drop a PDF here, or click to choose
Your file never leaves your device.
Burning in redactions…
Preparing pages…

Three PDF shapes confuse redaction conversations. Native text PDFs store real character operators — select, search, copy. Image-only scans are pictures per page; without OCR there is no text to search until someone adds it. OCR'd scans are the trap: the page looks like a fax photo, but a parallel text layer encodes what the OCR engine read — names, account numbers, diagnosis codes — aligned invisibly under the pixels.

Native versus scan without OCR is redact scanned PDF vs native PDF. Copy-paste failure as a test is can you copy-paste text out of a redacted PDF?. This page is the OCR sandwich specifically.

How the OCR layer is structured

OCR software renders the scan as a background image (or inline image XObject) and injects transparent or same-color text objects positioned over the glyphs it recognized. Viewers use that layer for selection rectangles and search indexes. A human sees only the scan; the file still contains Unicode strings. Redacting the visible scan with a comment rectangle does not delete those strings unless the tool explicitly removes text objects and rebuilds the page.

Partial OCR makes it worse: the layer can be wrong visually but right enough to leak a SSN, or correct under a region you thought was ‘just a smudge.’ Never trust appearance alone.

What must be touched for safe redaction on OCR'd PDFs

Safe means the sensitive string is not recoverable by select, search, copy, or structure-tree extraction in the file you will send. Options include: delete underlying text objects and fill the area; rasterize the entire page (or marked regions) so text becomes non-selectable pixels; or rebuild a new PDF from flattened images. HidePDF’s model is rasterize pages with boxes burned in — the export you download should fail search in marked areas if boxes cover the text completely.

Visual-only cover on the image layer alone leaves the ghost text. Sanitization tools that ‘remove hidden text’ without flattening may still miss misaligned OCR crumbs outside your box — always search the export.

How to test an OCR'd PDF before and after

STEP 01

Search the original

Find a distinctive string that must disappear. If search hits on a scan-looking page, you have OCR.

STEP 02

Try select-all paste

Copy page text into a plain editor. Invisible layer text may appear without looking selectable.

STEP 03

Redact with burn-in

Draw boxes with margin in the tool above. Download.

STEP 04

Search and paste again

Same strings should fail. One hit means the OCR layer or image still leaks.

Scenarios

A medical record scan OCR'd for filing. Patient name is searchable under the scan. Black highlight on the image is not enough.

A court exhibit from a copier ‘searchable PDF’ mode. Opposing counsel searches your production. OCR text must be destroyed, not hidden.

A legacy archive you re-OCR'd. Two text generations can coexist with different offsets — cover visually aligned to the image may miss misregistered text.

How to use HidePDF on OCR'd scans

  1. Open the PDF locally in the tool.
  2. Box every sensitive region on every page — OCR text can sit slightly offset from visible ink.
  3. Download. Search. Paste. Treat OCR like live text until tests pass.

More on scan types: redact scanned PDF vs native PDF. Paste test: can you copy-paste text out of a redacted PDF?

Double OCR and misregistration

Documents OCR'd twice with different engines can contain overlapping text layers offset by a few points — search finds strings in places that do not align with the visible scan. Box margins must be generous. Some production workflows OCR after redaction burn-in specifically to prevent ghost text — your process should match legal team guidance.

Search indexes built at ingest time (eDiscovery platforms) may retain pre-redaction text even after you replace the viewer PDF — another reason production teams regenerate load files, not only swap attachments.

Related guides

See also:

Frequently asked questions

How do I know if my PDF has an OCR text layer?

Try selecting text with the cursor even though the page looks like a scan. If words highlight and copy to Notepad, an invisible text layer exists — often from Adobe Acrobat OCR, ABBYY, or scanner software.

Does drawing a black rectangle in a viewer remove OCR text?

Usually not if the rectangle is an annotation or image stamp on top of the page. The hidden text objects remain in the structure tree. Search and copy may still work.

What does HidePDF do on OCR'd scans?

HidePDF rasterizes each page with your marked boxes burned into the page image for the export. That removes the usual copy-and-search path through live text in marked regions — verify by searching and paste-testing the download.

Does HidePDF send the OCR'd PDF to a server?

No. The PDF is processed in local browser memory on this page.