Redact a scanned PDF vs a native PDF
Understand the technical difference between image-based scans and text-layer PDFs — and redact each type so black boxes cannot be lifted away.
Verification is the step that catches redaction failures before a PDF leaves your control. A good check combines visual review, text search, copy-paste testing, a second PDF viewer, and a page-by-page scan for repeated identifiers in headers, footers, attachments, and OCR text.
HidePDF runs entirely in your browser, which matters for confirming a PDF is safe before you hit send. Your PDF never leaves your device; processing happens in local memory on your device. Server-based tools that only draw shapes often pass visual inspection yet fail copy-paste tests.
How HidePDF works
Open your draft redacted PDF
Load the export in the redaction tool on this page or your viewer, then follow the verification steps below on redaction verification.
Draw permanent black boxes
Click and drag over leftover selectable text and comment layers. Each box is burned into the rasterized page so the original text layer cannot be recovered.
Download and verify
Save the redacted PDF, then try select-all and search in a viewer. Redacted regions should not return readable text.
Scanned PDF vs native PDF redaction
A digitally native PDF usually carries a selectable text layer, so redaction must destroy or remove those characters—not just cover them. A scanned PDF is typically a flat page image: there is no text string to delete until OCR is added, so concealment has to rewrite the pixels themselves.
That is why an overlay-only black box can fail on a scan when another viewer strips the markup and reveals the original image underneath. The sections below walk through each type, how to tell them apart, and how to redact them correctly in HidePDF.
Related guides
Explore more ways to redact PDFs privately, or use the redaction tool above:
Frequently asked questions
Can I redact a scanned PDF the same way I redact a text-based PDF?
You use the same HidePDF canvas, but the mechanics differ underneath. On a native PDF, redaction must destroy or remove the text layer in the selected region. On a scan, there is no selectable text — you are covering or replacing image pixels. In both cases, export a flattened file and verify that nothing sensitive remains recoverable.
Why did my black box over a scanned SSN still show text when someone tried to lift it?
Some tools place a vector rectangle above the page instead of rewriting the page image. That overlay can be removed or ignored by another viewer, revealing the scan underneath. HidePDF burns redactions into a flattened export so covered pixels are part of the page content, not a removable sticker.
How do I tell whether my PDF is scanned or digitally native?
Try selecting text with your cursor or using find-in-document search for a word you can see. If text highlights and copies, you likely have a text layer (native or OCR'd). If nothing selects and search finds nothing, you are looking at a flat page image — treat it as a scanned PDF for redaction purposes.
Does HidePDF send my PDF to a server while I redact it?
No. The file is read into your browser, rendered locally, and exported from your device. HidePDF does not require an account, and the document does not need to leave your machine for redaction to run.
Should I ask a colleague to test?
A second pair of eyes helps on legal productions - have them search for known SSN fragments. Give them a checklist so the review is consistent. Do not send the file broadly until the first verification pass succeeds.
Do OCR tools resurrect text?
OCR on flattened redacted pages should not recover burned-out regions. Re-OCR only non-redacted areas if needed. If OCR recovers covered text, the visible redaction did not actually remove the source content.
What search terms should I use during verification?
Use exact names, account fragments, dates of birth, addresses, case numbers, SSNs, merchant names, and unique phrases from covered sections. Search with and without punctuation. Include alternate spellings or abbreviations if the PDF uses them.
What if verification finds one missed redaction after export?
Go back to the original working copy, add the missing redaction, and export a new file. Do not try to patch the exported PDF with a simple drawing tool. Repeat the full verification checklist on the new export.
Not every PDF holds information the same way. A digitally native PDF usually contains a real text layer — characters your computer can select, search, and copy. A scanned PDF is different: it is typically a photograph or flat image of paper pages wrapped in a PDF container. The letters look sharp on screen, but to the file format they are pixels, not strings. That distinction is easy to miss until redaction fails in a surprising way.
This page explains how native PDF text layers work, how scanned PDFs differ, why the difference matters for redaction specifically, and how to handle each type correctly with HidePDF in your browser. The goal is practical: fewer false assumptions about what a black box actually destroyed, and a clearer verification habit before a file is emailed, filed, or posted.
How native PDF text layers work
When software “prints” or exports a document to PDF — from a word processor, spreadsheet, billing system, or form generator — it often embeds fonts and a structured text stream. Readers can highlight a Social Security number, search for an email address, or copy a paragraph into another app. Accessibility tools can also read that layer aloud. From a redaction perspective, that text layer is both convenient and dangerous: convenient because tools can target strings precisely; dangerous because covering text visually without removing the underlying characters leaves the secret intact for anyone who selects, searches, or extracts text.
Proper text redaction on a native PDF therefore has to change the page content itself. Removing or replacing glyphs in the selected region, flattening annotations into irreversible page content, and stripping related form-field or metadata echoes are part of making concealment stick. A rectangle that only sits on top of the page as a movable annotation is not the same as destroying the characters underneath.
How scanned PDFs differ
A scanned PDF usually starts life as a multipage image: a phone photo of a form, a flatbed scan of a medical printout, or a fax converted to PDF. Optical character recognition (OCR) may be added later, which creates a text layer that sits over the image. Until OCR exists, there is nothing to “find” with Ctrl+F — the SSN is paint on a bitmap. Redacting that page means changing image pixels (or replacing the page image entirely), not deleting a string from a content stream.
Many people still draw a black box and assume the scan is safe. If the editing tool stored that box as a separate vector overlay, another person can delete the overlay, hide annotations, or open the file in a viewer that ignores the markup. The original scan pixels remain. That failure mode is specific to image-backed pages treated as if they were text documents.
Why this matters for redaction
Redaction is a permanence claim. Recipients, opposing counsel, journalists, or curious forwarders will test whether hidden content is truly gone. On native PDFs, the classic failure is leaving selectable text under a black rectangle. On scanned PDFs, the classic failure is leaving an intact page image under a removable overlay. Both look fine in a quick preview. Both can fail a one-minute adversarial check.
Mixed documents add confusion: a native cover letter attached to a scanned ID page, or a scan that later received OCR. One file can need text-layer destruction on page one and image-level coverage on page two. Guessing from the filename (“scan.pdf”) is unreliable; test the page in front of you.
How to redact each type correctly with HidePDF
HidePDF runs in your browser. You open the PDF on this page, mark regions, and export a flattened result from your device. Use the same careful marking habit for both PDF types, then verify with the method that matches the content:
- Native / text-layer pages. Select every sensitive string with enough margin to catch descenders and repeated headers. Export a flattened PDF so redaction is burned into page content rather than left as a loose annotation. After export, try selecting text in the covered area and searching for a known secret string — nothing should come back.
- Scanned / image pages. Treat the page as a picture. Cover the pixels that form the sensitive marks (SSN, DOB, signatures, barcodes). Export flattened output so the covered region is part of the page image. Then try common “lift the box” checks: open the export in a second viewer, confirm annotations are not independently toggleable, and make sure the underlying scan is not still extractable as an untouched image object.
- OCR'd scans. If search finds text on a page that also looks like a photograph of paper, you likely have image + text. Redact visually and verify both that pixels are covered and that search/copy no longer return the secret. An OCR layer can reintroduce the native-PDF failure mode on top of a scan.
In all cases, verify the file you plan to send — the export — not only the in-progress canvas. Keep the unredacted original separate and private.
How to tell which type of PDF you have
Selection test. Click and drag across a paragraph. If characters highlight, a text layer (native or OCR) is present. If you only get a marching selection box with no character highlight, you are likely on a flat scan.
Search test. Search for a distinctive word printed on the page. Hits imply text. No hits on clearly visible words imply image-only content (or a scan without OCR).
Zoom test. Extreme zoom on a scan often shows compression artifacts or paper texture; native vector text stays sharp longer. This hint is imperfect alone, but useful alongside selection and search.
Page-mix test. Check more than the first page. Portfolios and packet PDFs frequently mix exported statements with photographed IDs.
Practical takeaway
Native PDFs hide risk in the text layer. Scanned PDFs hide risk in the image pixels — and in tools that only overlay boxes instead of rewriting those pixels. HidePDF is built for browser-local, flattened redaction so you can address both failure modes without sending the document to a remote service. Identify the page type, mark thoroughly, export, then verify with selection, search, and a second viewer before the file leaves your control.