H
HidePDF
Redact PDFs in your browser. Nothing leaves your device.
100% local · fully local

In What Order Should You Redact, Merge, and Sign a PDF?

Redaction is rarely the only thing you do to a document. It gets assembled, OCR'd, stamped, compressed and signed too — and because redaction is destructive by design, it is an ordering constraint rather than a step you can slot in wherever it fits.

PDF
Drop a PDF here, or click to choose
Your file never leaves your device.
Burning in redactions…
Preparing pages…

Nearly every other redaction guide — including most of the ones on this site — answers a question about one step in isolation: did the black box actually remove the text, does the OCR layer survive, can the recipient copy-paste the covered words back out. Those questions assume redaction is the whole job and the file is otherwise finished.

In real document work it almost never is. A packet gets assembled from several files. A scan gets OCR'd so somebody can search it. Pages get Bates-numbered or headed with a case name. The result gets compressed to clear an email or portal size limit. Then it gets signed, notarized, hashed, or logged. Redaction is one step among six or seven, and the thing nobody tells you is that the sequence is not free.

It is not free because redaction is the only step in that list that is deliberately destructive. Merging adds. OCR adds. Stamping adds. Compression degrades but preserves structure. Redaction removes, permanently, and in a burn-in tool it removes more than the rectangle you drew — it replaces the entire page with a picture of itself. Any step that depends on the thing redaction destroys has to happen on the other side of it. Get the order wrong and you do not get an error message; you get a finished file that quietly lost something, or a redaction that was never binding in the first place.

The rule that decides nearly all of it

You can derive most of the correct order from one observation: a burn-in redaction converts a structured document into a sequence of flat images. Everything that was structure — selectable text, an OCR layer underneath a scan, link annotations, form fields, bookmarks, tagging for screen readers — stops existing. Everything that was appearance becomes permanent pixels.

That splits your other steps into three groups, and the groups tell you where they go:

  • Steps that consume structure go before. Anything that reads text out of the file, or edits text in place, or relies on form fields, has to run while that structure is still there. If you need to extract a table, pull a data set, or fill in a form, do it first.
  • Steps that produce structure go after. OCR, text stamping, tagging, re-adding links, building a table of contents. Run them before the redaction and they are flattened away; run them after and they survive into the file you send.
  • Steps that bind to the finished bytes go last. Digital signatures, notarization, checksums, manifests, and the entry you write in a redaction log. These are statements about a specific file, so they have to be made about the file you are actually sending — which is the post-redaction one.

A working order, step by step

  1. Decide the scope first, on paper. Which document, which pages, which categories of information. This costs nothing and it prevents the single most expensive mistake in the pipeline, which is discovering halfway through that a different file was the one that needed redacting.
  2. Do your extraction and data work now. If anyone needs figures, quotes, a copy of the text, or the contents of form fields out of this document, take them while the text still exists. After redaction there is nothing to extract from.
  3. Choose merge-then-redact or redact-then-merge deliberately. This is the real decision on the page, and it is covered in detail below.
  4. Redact, and verify immediately. Search the output, try to select text, try to copy-paste, and look at the file as a stranger would. Verify before you build anything on top of it — every later step makes a mistake more expensive to unwind.
  5. OCR, if the recipient needs search. On the redacted output, never before it. The resulting text layer is generated only from pixels you chose to publish, so there is no older layer underneath it to leak.
  6. Stamp, number, and head the pages. Either here, as live text, or back at step 3 as pixels. Both are defensible; the trade-off is below.
  7. Compress only if you must, and gently. A burn-in export is already a set of JPEG page images. Re-compressing JPEGs stacks artefacts on artefacts, and the first thing to go soft is small print — exactly the dense, low-contrast text in a footnote or a figure legend that a reader most needs to be legible.
  8. Sign, notarize, hash, log, send. All of these describe the final file. Do them in that order and nothing you do afterwards contradicts them.

Merge first, or redact first?

This one genuinely has no universal answer, so here is the trade-off stated plainly. A burn-in tool of this kind rasterizes every page of whatever file you hand it, not only the pages carrying boxes. That single fact drives the decision.

Merge first if the packet is small, or if most of it is sensitive anyway, or if consistency is the priority. You get one continuous page numbering to work against, you make one pass rather than five, and — importantly — you can keep your box sizing uniform across the whole packet, which matters because the width and count of your rectangles is a cross-document signal in its own right. The cost is that the entire packet comes back as images, including the parts that needed nothing done to them.

Redact components first if the packet is large and mostly clean. Flatten the three pages that need it, leave the other eighty-seven with their text layer, bookmarks and links intact, and merge afterwards. The recipient can still search most of the document. The cost is coordination: separate files, separate verification passes, and a real risk of merging the wrong version of a component — which is its own well-documented category of accident.

A reasonable default for mixed packets: redact the sensitive components individually, verify each one, merge, then verify the assembled packet again before it leaves your hands.

What this tool actually does, so you can plan around it

Ordering advice is only useful if it matches the tool's real behaviour, so here is what the code on this page does, read from the source rather than assumed:

  • One PDF at a time, PDFs only. The file gate checks the type and the .pdf extension and rejects everything else, and there is no batch mode. Merging and splitting are separate jobs for separate tools.
  • Every page is rasterized, not just marked ones. The export loops over all pages, renders each at twice the page's natural size — roughly 144 DPI, since PDF points are 72 to the inch — and embeds each as a JPEG at quality 0.92.
  • The output is a brand-new document. The code creates an empty PDF and adds page images to it. Nothing from the original object tree is copied, which is why there are no annotations, no form fields, no bookmarks, no link targets and no signature in the result. They are not stripped so much as never carried over.
  • No text, no detection. There is no text search, no pattern matching and no automatic finding of names or numbers. You draw the rectangles, and the tool draws exactly the rectangle you drag.
  • Document properties are rewritten. The export sets the PDF title to "Redacted Document" and the producer and creator to "HidePDF", so the original application and title do not ride along.
  • The filename does. The download is named after your original file with -redacted appended. If the filename itself is sensitive — a client name, a case number, a diagnosis — rename it, either before you load it or after you save it. Renaming is a step in the pipeline too, and it is the one people forget because it happens outside the document.
  • All of it runs in the browser. The processing happens in local memory on this page, with no upload required and no account.

Common mistakes and misconceptions

Signing first and redacting after. The most consequential ordering error there is. With a rebuild-style tool the signature is not invalidated, it is absent — you have produced an unsigned document that looks signed, because the visible signature graphic is now part of the page image. If the signed original mattered, you now need a fresh execution, not a quietly edited file.

Paying for OCR before redacting. The text layer does not survive, so the money and the processing time are spent twice. OCR last.

Treating a valid signature as a redaction check. These answer different questions, and signature validity is a weaker guarantee about appearance than it sounds — the NDSS 2021 Shadow Attacks work demonstrated altering what a signed PDF displays while the signature still validated. Verify the redaction directly.

Compressing at the end to hit a size limit. Understandable, but the redacted export is already image-based and therefore already large and already lossy. If size is a constraint, plan for it earlier — split the packet, or redact components so most pages stay as compact text.

Assuming flattening is reversible if you kept the original. It is not reversible in the file; you are relying on holding a second, unredacted copy. That copy is now the most sensitive object in the job, and keeping two near-identical files straight is a known failure mode.

Writing the redaction log before producing the final file. A log entry that describes a draft is worse than no log, because it is confidently wrong. Log the file you send, with whatever identifier you can actually tie to it.

Verifying once, at the end. Verify straight after redaction and again on the final packet. Steps between the two — merging, stamping, compressing, converting — are exactly where a correct redaction gets replaced by an older version of the same document.

Related guides

See also, or use the redaction tool above:

Frequently asked questions

Should I redact before or after merging documents into one packet?

It depends on how much of the packet you are willing to turn into page images. A burn-in redaction tool of this kind rasterizes every page of the file it is given, not only the pages you drew boxes on. So if you merge a 90-page packet and then redact three pages of it, all 90 pages come back as images: no selectable text, no bookmarks, no links anywhere in the packet. If you redact the three-page component on its own and merge afterwards, only those three pages are flattened and the other 87 keep their text layer. The argument for merging first is consistency: one continuous page numbering, one pass, and one chance to size your boxes the same way throughout, which matters because box geometry is a cross-document signal. The usual compromise is to merge first if the packet is small or uniformly sensitive, and to redact components first if most of the packet is clean and the recipient needs to search it.

Can I OCR a PDF first and still have it searchable after redaction?

No, and this is the most common ordering mistake. OCR adds an invisible text layer underneath the page image. A burn-in redaction rebuilds every page as a fresh picture and assembles those pictures into a new document, so that text layer is not edited or filtered — it is simply not carried across. The output has no selectable text at all, which is exactly why the covered words cannot be copied or searched out of it, but it also means any OCR you paid for beforehand is gone. If the finished file needs to be searchable, OCR the redacted output as the last content step, after you have verified the redaction. That way the only text in the file is text generated from pixels you were willing to publish, and there is no older layer hiding underneath it.

If I redact a signed PDF, does the signature break or just disappear?

With a rebuild-style tool it disappears rather than breaks. The export creates an empty PDF document and adds page images to it; nothing from the original file's object structure is copied over, so the signature dictionary and its certificate are not present in the output to be checked at all. A validator does not report a tampered signature, it reports an unsigned file. Either way the practical answer is the same: redact first, sign last. It is also worth knowing that a valid signature is weaker evidence of an unaltered appearance than people assume — research presented at NDSS 2021 as Shadow Attacks: Hiding and Replacing Content in Signed PDFs, by Christian Mainka and colleagues, demonstrated changing what a signed PDF displays without invalidating its signature. So do not treat signature validity as your redaction check. Verify the redaction on its own terms, then sign.

Where do Bates numbers and page stamps belong in the order?

Both positions work, and they produce genuinely different files, so pick deliberately. Stamp before redacting and the numbers are rendered into the page image along with everything else: permanent, visually faithful, impossible to shift or strip, but no longer machine-readable and no longer editable if you later need to renumber. Stamp after redacting and the numbers stay as real text or annotations, so they can be extracted by a review platform or an e-filing system, and they can be corrected — but they sit in a layer that someone downstream can also remove. If the numbering is how other people will cite the document in correspondence or a filing, stamp before and accept the permanence. If the numbering is a working artefact in a review workflow, stamp after. What you should not do is stamp, redact, and then assume the stamps are still text: check by trying to select one in the finished file.