Can a PDF's Filename Leak What You Redacted?
Every other redaction question is about the inside of the document. The name on the outside is a separate disclosure, it is not covered by any black box, and it travels to places the document itself never goes.
Almost every redaction question is a question about the interior of a document. Did the black rectangle remove the text underneath it, or just sit on top of it? Does an OCR layer survive? Can the recipient copy and paste the covered words back out? Is there anything in the metadata, the annotations, the revision history? Those are all questions about bytes inside the file, and they all have the same shape: you examine the document, and either the information is in there or it is not.
The filename is a different kind of object entirely. It is not inside the document — it is a label attached to the document by the filesystem, and later by the email client, the archive, the share link and the recipient's filing system. No redaction tool in existence touches it, because it is not part of what the tool was handed. You can produce a technically flawless redaction, verify it three ways, and still deliver the withheld fact in the attachment chip above it, in nine-point grey type, to everyone on the thread.
This is a routine failure, and it is routine for a structural reason: the filename is written at the beginning of the job, when you are thinking about finding the file again, and it is read at the end of the job, by someone you had not thought about yet. Hendricks-termination-harassment-claim-DRAFT.pdf is a perfectly sensible name for a working file on your own machine. It is a disclosure when it arrives in opposing counsel's inbox with every relevant passage correctly blacked out.
Why the name is its own disclosure surface
Three properties make a filename behave differently from anything else on this site, and they are worth stating separately because the mitigations differ.
- It is outside the scope of every verification check you know. The standard checks — search the output for a covered term, try to select text, look at the document properties, inspect the object tree — all operate on the file's contents. A filename passes all of them trivially, because it is not in there to be found. The one check that catches it is looking at the name, which is exactly the thing your eye stops registering after the fortieth time you have opened the file.
- It is copied, not moved. When a document is delivered, the bytes go to one place but the name is transcribed into several: a message header, a log line, a database row, a folder listing, a notification email. Deleting or replacing the document does not retract those copies. This is the property that makes a filename disclosure hard to clean up after, and it is why the only real fix is upstream.
- It is often read when the document is not. People scan attachment names, subject lines and folder listings constantly and open documents rarely. A revealing name is therefore more likely to be read than the carefully redacted page behind it — including by people who were only ever meant to handle the message, not its contents: an assistant, a mailroom, an intake clerk, whoever manages the shared drive.
There is a related point worth making precisely, because it cuts the other way. In HTTP, the name a browser saves a download as does not have to come from any file on a disk at all: RFC 6266, Use of the Content-Disposition Header Field in the HTTP Protocol (Reschke, 2011), defines the filename parameter a server uses to suggest one. The suggested name travels in a header, separate from the document's bytes. If you publish a document through a portal or a CMS that sets that name from a database record rather than from your upload, your careful rename may be discarded — or a name you never chose may be applied. Worth checking once per system rather than assuming.
What this tool does with your filename, read from the source
Advice about naming is only useful if it matches what the software actually does, so here is the behaviour of the code on this page, read from the source rather than assumed:
- The download inherits your name. On Download, the code takes the name of the file you loaded, strips a trailing
.pdf, and appends-redacted.pdf. SoHendricks-harassment-claim.pdfcomes back asHendricks-harassment-claim-redacted.pdf. Every word you were worried about is still in the name, now with a suffix that advertises there is something to be curious about. - The document's own properties are scrubbed, which is the confusing part. The export sets the PDF title to "Redacted Document" and the producer and creator to "HidePDF". Because the output is built as a new document with nothing copied from the original object tree, no source title, no authoring application and no annotations carry across. So the inside of the file is clean and the label on it is not — which is precisely the asymmetry that makes this easy to miss. If you check the document properties and see them scrubbed, you can come away with an impression of thoroughness that the filename does not deserve.
- Nothing about the name is transmitted anywhere. The loading, the rendering and the export all happen in your browser's memory on this page, with no upload required and no account. The name shown in the toolbar is read from your own file picker and stays on your device.
- The rename is yours to make, at either end. Rename the source file before you load it and the export is named correctly automatically. Rename the download afterwards and you get the same result, with one more chance to forget. The first option is better for exactly that reason.
Where the name goes that the document does not
This is the part that makes a filename disclosure expensive. Work through the list for whatever channel you actually use.
- The attachment chip and the message body. Visible to every recipient and every person later added to the thread, including on forwards where the document itself may be stripped by a gateway but the name survives in the quoted message.
- Subject lines and your own covering note. People reflexively paste the filename into the subject or write "attached is the Hendricks harassment claim". Redacting the document and then describing it accurately in the email is a very common own goal.
- Mail logs, archives and retention systems. Attachment names are routinely indexed by mail servers, archiving appliances and e-discovery collection tools. That index outlives the message, is searchable by people other than the recipient, and is frequently itself discoverable.
- The index of a zip archive. Covered in the FAQ below, but the short version: zipping with a password does not normally hide the names of the files inside.
- Share links and their notification emails. A cloud share is usually named for the file, and the automated "X shared a document with you" message carries that name to the recipient's inbox — and often to a notification in a chat tool, where it sits in a channel's searchable history.
- Portal receipts, confirmation pages and filing indexes. Submission systems commonly echo the uploaded filename back on a receipt, in a confirmation email and in a docket or case index. Those are frequently more public than the document, and entirely outside your control once submitted.
- The recipient's own system. Your name becomes a row in their document management system, often with the file's text indexed alongside it. If their index is ever exported, reviewed or produced, the name goes with it.
- Screens other people can see. Downloads bars, recent-file lists, tab titles and desktop icons. A filename is the single most likely thing to be read off your screen during a call or a screen share, because it is short and it is in the chrome rather than in the document.
A naming convention for a file that leaves
You do not need a policy document for this. You need a rule you can apply in the four seconds before you attach something.
- Name the role, not the contents. What the document is in this transaction — exhibit, schedule, statement, application — rather than what it says. Roles are safe to publish because the recipient already knows them; contents are the thing you spent the last twenty minutes covering up.
- Mark the released version at the front of the name. Put
redacted-orrelease-first, where it cannot be truncated away in a narrow column, an attachment chip or a mobile picker. The marker does double duty: it tells the recipient what they have, and it tells you which of your two copies you are about to send. - Strip the four categories that cause actual harm. Personal names, identifying numbers (case, account, claim, policy, member, employee), conditions and diagnoses, and amounts. If a word in the filename falls into one of those and you redacted the same word inside the document, you have contradicted yourself.
- Drop your working vocabulary.
DRAFT,v7,FINAL-final,per-Sarah,do-not-send,privileged,strategy. These describe your internal process and your internal assessment of the document, which is not information the recipient was given and occasionally not information you would want to defend. - Add a date in a sortable form.
2026-10-03rather thanOct3. It costs nothing, it disambiguates successive releases of the same document, and it is the single most useful thing you can give a recipient who has to find the file again in six months. - Rename before you load, then say it out loud once. Read the final name back as though you were the recipient reading it in their inbox. That is the whole check, and it takes a second.
A workable result looks like redacted-exhibit-c-2026-10-03.pdf: unambiguous, sortable, boring, and it reveals nothing that the black boxes were put there to hide.
Common mistakes and misconceptions
Believing a clean metadata panel means a clean file. The two are unrelated. With the export from this page the document properties are rewritten and nothing from the original is copied, so the properties look immaculate while the download keeps your original name. Checking one tells you nothing about the other.
Appending the marker instead of leading with it. very-long-descriptive-case-name-redacted.pdf loses the only part that mattered as soon as the display width runs out, which it does in most attachment chips and almost every mobile client.
Assuming a zip with a password conceals the contents list. It normally does not. Ordinary archive encryption protects the file data and leaves the index readable.
Renaming only the copy you send, and only sometimes. If the rename is a thing you remember to do, you will eventually not remember. Make it part of producing the export — rename the source, then redact — so the correct name is the default rather than a discipline.
Over-correcting into scan001.pdf. Anonymous names are how the unredacted original gets sent instead of the redacted copy. The name is also a safety control for you, and a name that distinguishes nothing cannot perform that job.
Forgetting the folder and the archive around it. A clean filename inside Litigation/Hendricks-v-Nortek/ discloses the same facts the moment the folder is shared, synced or zipped with its path intact. Check the container, not just the leaf.
Writing the revealing name into the covering email anyway. The most common version of this mistake. If the filename needed sanitising, so does the sentence describing it.
Treating a revealing name as harmless because the box is correct. It is the same category of error as leaving selectable text under a rectangle: the information is available to the recipient without any effort or ingenuity on their part. That the mechanism is mundane does not make the disclosure smaller.
Related guides
See also, or use the redaction tool above:
- Keeping the redacted and unredacted versions of a PDF straight
- Can a PDF's author or editing metadata reveal who redacted it?
Frequently asked questions
If I rename the file, is the old name still stored inside the PDF?
A PDF has no standard field that records the name of the file on disk, so renaming the file in your operating system is usually enough on its own. The qualification is that some producing applications copy a source title or a source path into the document's own properties — the Title entry, or XMP media-management fields that can hold a reference to the file a document was derived from — and renaming the file afterwards does not touch those. That is a property of whichever program made the original, not of the name you typed. With the export from the tool on this page the question is settled a different way: the code builds a brand-new document and copies nothing from the original object tree, then sets the title to "Redacted Document" and the producer and creator to "HidePDF". So no source title, no source path and no original application ride along in the output. The download name is the one thing that does carry over, and that is yours to change.
Does putting the files in a password-protected zip hide the filenames?
Usually not, and this catches people out because it feels like it should. The ZIP format keeps an index of the archive's contents — the central directory — separately from the compressed file data. Ordinary password protection encrypts the data, not that index. The .ZIP File Format Specification says so explicitly: it notes that information leakage can occur even though a file is stored encrypted, through exposure of data such as a file's name, its original size, timestamp and CRC32 value, and it introduced a distinct Central Directory Encryption feature in version 6.2 specifically to conceal that metadata. That feature is opt-in, is not what the zip command in a typical file manager produces, and programs that cannot read an encrypted central directory cannot open the archive at all. So assume a recipient, a mail gateway or anyone who intercepts the archive can list the names inside it without the password. If the names matter, fix the names.
Should I just rename everything to something meaningless like document1.pdf?
No — that trades one problem for two worse ones. A meaningless name makes it very hard for you to tell the outbound copy from the unredacted original you still hold, and mixing up a near-identical pair is a far more common and more damaging accident than a revealing filename. It also pushes work onto the recipient, who now has to open every attachment to find out what it is, and who will rename or mis-file it in ways you cannot predict or cite later. Aim for neutral and specific rather than empty: name the file for the role it plays in the transaction, not for its contents. Something like redacted-exhibit-c-2026-10-03.pdf tells everyone which document this is and that it is the released version, without naming a person, a condition, an amount or a case theory.
The recipient already has the revealing filename. Is there anything useful to do?
You cannot retract a name that has already been delivered, and resending the same document under a cleaner name does not remove the first one — it adds a second copy and draws attention to the difference. What is worth doing is bounded and practical. Work out where the name actually landed: the message itself, any forward of it, a shared-drive folder, a portal receipt, an index in the recipient's document system. If the disclosure is material — a name, a case number, a diagnosis, an amount you were specifically obliged to withhold — tell the recipient plainly and ask them to delete the copies rather than quietly hoping. If there is an obligation attached to the withheld material, the disclosure of the name is the thing you report, not the file. Then change the step that caused it: rename at the point of export, before the file ever reaches a compose window.