Why Redacting Names Isn't Enough: Indirect Identifiers in PDFs
A page with a black box over every name can still describe exactly one person. This is not a question about whether your redaction held — it is a question about whether you covered the right things in the first place.
Nearly every redaction guide — including most of the ones on this site — answers a mechanical question. Did the black rectangle actually destroy the text underneath, or is it sitting on top of a selectable layer? Can OCR pull the words back? Does the metadata still name the author? Those are questions about whether a redaction held. They all assume you already know which words needed covering.
This page is about the assumption itself. It is entirely possible to run a flawless permanent redaction — boxes burned in, text layer gone, copy-paste dead — and hand over a document that still tells a reader precisely who it is about. Nothing failed technically. The failure happened earlier, when someone decided that redaction meant "cover the names and the ID numbers" and stopped there. If your document is going to a party who knows something about the context it came from, that is not a safe stopping point.
Direct identifiers versus indirect identifiers
A direct identifier points at one person on its own. A full name, a Social Security number, an account number, a passport number, an email address, a phone number, a signature. Everybody redacts these. They are what the word "redaction" brings to mind, and they are the easy half of the job.
An indirect identifier — sometimes called a quasi-identifier — does not single anybody out by itself, but narrows the field. A job title. A department. A hire date. A date of birth. A postal code. A hospital. A court. A vehicle model. The fact that someone was the only person on a particular shift. Each one describes a group. The problem is that groups intersect, and intersections shrink fast. Two or three coarse details that each describe thousands of people can, in combination, describe one.
This is the reasoning behind formal de-identification standards rather than a novel idea. HIPAA's Safe Harbor method, for example, is written as an enumerated list of identifier categories that must be stripped, and that list deliberately reaches past names into things like dates and geographic subdivisions precisely because those combine. Data-protection law in the EU likewise defines personal data broadly enough to cover information from which a person can be identified indirectly. If a specific standard applies to your document, read its current text — the rule is the authority, and the categories and thresholds it sets are the ones that count, not a summary on a tool's website.
The combination test
Here is a method you can actually run on a document in a few minutes, without any statistical training.
- Define the reader. Write down, in one sentence, who will see this file and what they already know. "Opposing counsel, who has the full employee roster" is a completely different reader from "a member of the public who found this on a website." Re-identification risk is never a property of the document alone; it is a property of the document plus the reader's outside knowledge.
- List every remaining descriptive detail. Walk the document and note each fact about the person that survives your obvious redactions: role, dates, locations, amounts, sequence numbers, relationships, unusual events.
- Ask "how many people?" for each one alone. Roughly, in the population your reader can see, how many people does this detail fit? "Employee" fits everyone. "Night-shift warehouse supervisor" may fit four.
- Ask it again for pairs and triples. This is the step people skip. "Night-shift warehouse supervisor" plus "started in March" plus "the incident on the 14th" may fit exactly one person, even though no single element did.
- Break the smallest combination. You do not have to cover all three. Covering one element of a narrowing combination is enough to widen the group again — and it usually preserves more of the document's usefulness than blanket redaction does.
- Re-run the test on what is left. Removing one element sometimes makes another combination load-bearing. One pass is rarely enough on a document of any length.
The test has a useful side effect: it tells you when you are over-redacting. A detail that fits thousands of people and does not combine with anything else on the page is usually safe to leave, and leaving it is what keeps the document worth reading.
Where indirect identifiers hide in a PDF
Body text is the obvious place to look, and the one people actually check. These are the places that get missed:
- Headers, footers, and Bates or control numbers. A sequence number can tie a document to a position in a production set, which can tie it to a custodian.
- Letterheads, stamps, and logos. A clinic logo is a location. A regional office stamp is a geographic narrowing you did not intend to publish.
- Handwriting and initials in margins. Initials are often treated as already-anonymous. Within a small organisation they are not.
- Dates other than the one you were thinking about. Signature dates, fax banners, received stamps, appointment times. Timestamps combine viciously with everything else.
- Amounts that are unusual. A salary, a settlement figure, or an invoice total that is distinctive functions as an identifier to anyone who already knows the number.
- Charts and images. A photo, a floor plan, a scanned badge, or a graph with one obvious outlier can carry identity that no text search would ever surface.
- The file itself. Filenames and document properties travel with the PDF and often still carry a person's name long after the pages are clean.
What this tool does and does not do
Worth being explicit, because it shapes how you should work. HidePDF has no text search, no pattern matching, and no automatic detection of names, numbers, or identifiers of any kind. It takes a PDF, lets you draw black boxes on each page by hand, and rebuilds the pages so the covered content is permanently gone rather than hidden under a shape. All of that happens in this browser tab — your file never leaves your device.
The consequence is that every judgment described on this page is yours to make. There is no automated safety net that will flag a date you left in, and no tool on the market would reliably catch a combination-based identifier anyway, because the combination only becomes identifying in light of what your particular reader knows. That is an argument for writing your decisions down. A short checklist for the document type — which fields get covered, which get generalised, which stay — applied the same way to every file beats a fresh judgment call each time, especially when you are working through a stack. Consistency across a set of documents is its own failure mode: redacting a name in one file and leaving it in another lets a reader join the two.
Common mistakes and misconceptions
"I redacted the name, so it is anonymous." Anonymity is about whether a reader can work out who this is, not about whether a particular field is blacked out. Judge the result, not the checklist item.
Treating initials, employee numbers, or pseudonyms as redaction. Replacing "Jane Doe" with "Employee 4" does not help if the recipient can map employee numbers, or if the rest of the page describes Employee 4 well enough to name them.
Forgetting that the reader has other documents. The most common re-identification route is not clever inference, it is a second file. Your redacted version plus an earlier unredacted version, a public filing, or a directory is often all it takes.
Generalising with a black box. Covering an exact date does not turn it into a year — it removes it entirely, which may make the document useless. If you want a coarser value rather than no value, change it in the source document before you produce the PDF. A redaction tool can only remove what is on the page; it cannot rewrite it.
Assuming small numbers are safe because they are numbers. A count of one in a table cell ("1 employee in this category") is an identifier dressed as a statistic. Small cells in cross-tabulated tables are a classic disclosure route.
Over-redacting to avoid thinking. Blacking out most of a page is not caution, it is a different kind of failure — documents get rejected, resubmitted, and handled twice, and each extra handling is another chance for the original to go somewhere it should not.
Related guides
Explore more ways to redact PDFs privately, or use the redaction tool above:
- Hide Personal Information in PDF
- Redact the same name or number consistently across multiple PDFs
- Over-redacting a PDF: when covering too much becomes the problem
- Redacting Numbers: When the Totals Give Them Away
Frequently asked questions
If I have covered every name and ID number, is the document de-identified?
Not necessarily. Names and ID numbers are direct identifiers, but a document can still point to exactly one person through a combination of details that are individually harmless — a job title, a date, a location, a rare circumstance. The right question is not whether you removed the obvious fields, but whether anything left on the page narrows the population down to one.
What is the combination test for deciding what to redact?
For each remaining detail, ask how many people it could plausibly describe in the population your reader can see. Then ask the same question about pairs and triples of those details together. If any combination narrows the group to one person, or to a group small enough that a reader could guess, that combination has to be broken by covering at least one of its parts.
Does HidePDF find sensitive information for me automatically?
No. HidePDF has no text search, no pattern matching, and no automatic detection of names or numbers. You draw black boxes on the page yourself, and the tool burns them permanently into a rebuilt PDF. That means the judgment about what counts as identifying is entirely yours — which is exactly why a written checklist beats redacting from memory.
Can generalising a value be safer than blacking it out?
Often, yes, but you have to do it in the source document before you export the PDF. Replacing an exact date of birth with a year, or a street address with a region, keeps the document useful while widening the group each value describes. A redaction tool can only cover what is already on the page; it cannot rewrite a value into a coarser one.