Can the Size of a Redaction Box Reveal What's Underneath?
A black box is not blank. Its width, its height, how many there are and where they sit are all things you published, and all of them were decided by the thing you were trying to hide.
Almost every redaction warning you will read — here and elsewhere — is about something that survived underneath. Live text sitting below a drawn rectangle. An OCR layer below a scanned image. A form field value, a comment, a revision, an author name in the document properties. The common shape of those problems is that the tool did not actually remove what you asked it to remove, and the fix is a better tool or a better export.
This page is about the opposite situation. Assume the removal worked perfectly. Assume the page has been rasterized, the covered pixels are irretrievably gone, and no amount of copy-pasting, OCR or file forensics will bring back a single character. There is still something on that page that you published and that carries information about the secret: the rectangle itself.
A black box has a width, a height, a position, and a count. You did not pick any of those arbitrarily — you picked them by dragging around the thing you were hiding. So the box is, unavoidably, a measurement of the hidden content. That measurement does not degrade with flattening, because flattening is what makes it permanent. It is the one part of the redaction that is supposed to be visible.
What a rectangle actually tells a reader
Work through it the way someone reading your document would:
- Length. A box drawn tightly around a name is roughly as wide as that name was. In a proportional font the mapping is not exact, but it is a strong constraint: a box five characters wide rules out most of the candidate names you were worried about, and a box twenty-four characters wide rules out the short ones. Narrow the candidate set enough and the redaction has not hidden anything, it has just made the reader do one step of work.
- Count. Three separate boxes in a paragraph say three items were removed. Where the removed items are people, that is a headcount. Where they are line items in a schedule, it is a transaction count. Readers get this for free, from a glance, with no tooling at all.
- Shape and height. A box one line tall over an address block says the covered address was one line. A box that spans four lines says something quite different. A tall thin box in a table column says a figure, not a sentence.
- Field position. On a form, the position of the box is the field label, and the label is often printed right next to it. Covering the value in a box marked "Date of birth" conceals the date and announces that you have one.
- Internal consistency. This is the one people miss. If the same person is redacted in forty places and you drew each box tightly, all forty boxes are about the same width — and a different person's redactions are a different width. That lets a reader cluster the redactions, count how many distinct entities there are, and follow one of them through the document, which is frequently the thing the redaction was meant to prevent.
- What you did not cover. The boundary of the box is a statement that everything outside it was fine to publish. Reading the words on either side of a tight redaction tells you what grammatical role the hidden token played, and often its type: a surname, a dollar amount, a date, a city.
The research version of this problem
There is a formal, peer-reviewed version of the layout-leak argument, and it is worth knowing about because it is much sharper than the eyeball version above. In Story Beyond the Eye: Glyph Positions Break PDF Text Redaction (arXiv 2206.02285; published in Proceedings on Privacy Enhancing Technologies, 2023), Maxwell Bland, Anushya Iyer and Kirill Levchenko report that many PDF redactions are insecure because character positioning information is not removed along with the characters. Their stated result is that sub-pixel horizontal shifts in the positions of redacted and non-redacted characters can be recovered and used to deredact first and last names. The paper's abstract describes a vulnerability assessment of 11 popular PDF redaction tools, including Adobe Acrobat, finding that all of them leaked information about redacted text, and the successful deredaction of hundreds of real-world redactions, among them OIG investigation reports and FOIA responses.
Two qualifications matter for how you apply this. First, that specific attack operates on the text-layout data inside the PDF, so it targets files where the text objects and their positioning survive the redaction — it is not a claim that any flat image of a page can be reversed. Second, the attack is about precision: measuring at sub-pixel accuracy narrows a name to a short list where eyeballing a box width would only narrow it to a long one. The underlying insight is the same one on this page, and it does not go away when you rasterize. Geometry constrains content. A permanent, pixel-perfect redaction that is sized to its contents is still publishing a measurement; it just publishes a coarser one.
How to draw boxes that don't talk
The fix is not a setting. It is a habit about where you let the edges of the rectangle fall.
- Size to the container, not the content. Let the layout decide the box, not the secret. Extend to the end of the line, across the full width of the table cell or column, or over the entire form field including its empty space. The edges are then a fact about the document's design, which the reader can already see, rather than a fact about what you removed.
- Standardise within a category. Every redacted name gets the same width; every redacted amount gets the same width. If a long name and a short name produce identical rectangles, width-matching across the document stops working, and so does counting distinct entities by box size.
- Merge adjacent redactions. Where several items in a row or a list all need covering, one rectangle over the group hides the count. A column of five small boxes is a published number five.
- Think about the whole document, not the page. Box geometry is a cross-page signal by nature. Decide your sizing convention before you start, because going back to normalise forty boxes on page thirty is a much worse job than deciding on page one.
- Read the surviving sentence out loud. If "payment was made to ███ of ███, Ohio" leaves a reader with a one-line box after "of" and a state, you have hidden a city name and published its length and its state. Ask whether the surrounding words need to go too.
- Accept the trade-off deliberately. Every item above hides more than the minimum, and that has a cost — in readability, and sometimes in compliance with an order or a records rule that tells you exactly what must be disclosed. Where a rule governs, the rule wins. Where it does not, decide consciously rather than defaulting to the tightest possible box because it looks tidier.
- Check the finished file as a stranger would. Open the exported PDF, read a page you did not just work on, and ask what you could guess about each box from its size, its neighbours and its twins elsewhere in the document.
What this tool does with box geometry
HidePDF draws exactly the rectangle you drag. The code takes the start and end points of your drag, takes the absolute difference for width and height, and stores that rectangle for the current page. There is no snapping to word or line boundaries, no padding around the edges, no auto-sizing, and no detection of what sits under the cursor — the tool never reads, searches or interprets your document's contents at all. The only size rule in the source is that a drag under a few pixels in either dimension is thrown away as a stray click rather than kept as a sliver.
That is a deliberate design, and it means the entire subject of this page is in your hands. On download, each page is rendered to an image at twice display resolution, your rectangles are filled in solid black onto that image, and the images are assembled into a new PDF. The pixels under the black are gone for good. The black is not: it is now a permanent, unambiguous, precisely measurable feature of the published page, sized exactly as you dragged it. Permanence protects the content and preserves the geometry. Both of those are working as intended; only one of them is doing what you probably assumed.
Common mistakes and misconceptions
"It's permanent, so it's safe." Permanence answers the question of whether the covered pixels can be recovered. It does not answer the question of what can be inferred from the cover. The two are independent, and a flawless redaction can still be an informative one.
Drawing tightly because it looks more professional. A snug box around each name reads as careful work, and it is the version that leaks most. Tidiness and disclosure pull in opposite directions here.
Treating each box as an isolated decision. The strongest signal in box geometry is comparative — this box versus that box, on page 4 versus page 31. You cannot see it while working on one page at a time.
Assuming a proportional font saves you. Variable character widths make the mapping from box width to string fuzzy, not useless. Fuzzy is enough when the candidate list is already short, which it usually is by the time someone is measuring boxes.
Forgetting that the label survives. Covering a value on a form leaves the printed field name next to it, so the box announces the type of what you hid even when it conceals the value perfectly.
Over-correcting into an unreadable document. The advice here points towards bigger boxes, and bigger boxes have their own failure mode — removing context the recipient legitimately needed, or withholding material you were obliged to produce. This is a balance to strike, not a direction to run in.
Expecting the software to handle it. No redaction tool can standardise box widths for you unless it knows what a name is and where the names are. A manual, hand-drawn tool certainly cannot, and will faithfully reproduce whatever shape you gave it.
Related guides
See also, or use the redaction tool above:
- Why Redacting Names Isn't Enough: Indirect Identifiers in PDFs
- Redacting Numbers: When the Totals Give Them Away
Frequently asked questions
If the pixels are permanently gone, how can the box still leak anything?
Because the box is itself a piece of published information. You chose where to put it and how big to make it, and both of those choices were determined by the thing you were hiding. A rectangle that is 14 characters wide tells a reader the covered string was roughly 14 characters wide. A rectangle one line tall tells them it was one line. Two boxes in a paragraph tell them two things were removed, not one. None of that requires recovering a single pixel or attacking the file — it is read straight off the finished page by eye. This is a different category of problem from a recoverable redaction. A recoverable redaction is a failure of the tool. A talking box is a failure of the drawing, and it survives any amount of flattening, rasterizing or burning in.
Is there real research showing redacted text can be recovered from layout?
Yes, for one specific and well-documented variant. In "Story Beyond the Eye: Glyph Positions Break PDF Text Redaction" (arXiv 2206.02285, published in Proceedings on Privacy Enhancing Technologies 2023), Maxwell Bland, Anushya Iyer and Kirill Levchenko showed that character positioning information left behind in PDFs can be used to recover redacted first and last names. Their abstract reports a vulnerability assessment of 11 popular PDF redaction tools, including Adobe Acrobat, all of which leaked information about redacted text, and the deredaction of hundreds of real-world redactions including OIG investigation reports and FOIA responses. Their precise attack depends on sub-pixel horizontal shifts in the surviving character positions, so it applies to PDFs where text objects remain and is not a claim about flat scanned images. The coarse principle, though, carries over to any page a human can look at: the geometry around and inside a redaction constrains what the redaction can have contained.
What should I actually do differently when I draw the box?
Stop sizing the box to the text and start sizing it to a unit the document already uses. Extend a redaction to the end of the line, to the full width of a table cell or column, or to the whole field in a form, so that the edges of the box are determined by the layout rather than by the length of what you removed. Make every redaction of the same kind the same size, so that a short name and a long name produce identical rectangles and cannot be told apart or matched across pages. Where you are covering several adjacent items, consider one box over all of them instead of a row of small ones, which otherwise publishes an exact count. The judgement to watch is that each of these moves hides slightly more than the minimum, so weigh them against your disclosure obligations and against the recipient's need to read the rest of the page.
Does HidePDF adjust or standardise my boxes for me?
No. HidePDF draws exactly the rectangle you drag and nothing else. There is no snapping to words, lines or table cells, no padding added to the edges, no detection of text, and no auto-sizing — the tool has no idea what is underneath your cursor, because it never reads the contents of your document. The only sizing rule in the code is that a drag smaller than a few pixels in either direction is discarded as an accidental click. That means the geometry described on this page is entirely yours to control, and entirely your responsibility: if you drag tightly around a word, a tight box is what gets published. On download, every page is rebuilt as an image with your boxes burned in as solid black, so the covered pixels are genuinely gone — but the rectangle you drew is still visible, at the size you drew it. All of this happens in your browser, with no upload required.