H
HidePDF
Redact PDFs in your browser. Nothing leaves your device.
100% local · fully local

What's Actually Inside a PDF — and What Redaction Changes

A PDF is not a picture of a document. It is a collection of objects, only some of which produce the page you look at. This is the map: what each part holds, what it can give away, and which parts a black box on the page actually changes.

PDF
Drop a PDF here, or click to choose
Your file never leaves your device.
Burning in redactions…
Preparing pages…

Most of the guides on this site answer one question about one hiding place. Can text be copied out from under a box. Can OCR recover a flattened page. Do bookmarks carry names. Does the filename give it away. Each of those pages is useful, and each of them exists because the answer surprised somebody. But they are all symptoms of a single underlying fact about the file format, and nobody ever writes that fact down.

This page is the other kind of page. It is not a scenario and not a yes-or-no question — it is the anatomy. Read it once and you stop needing a guide for each new hiding place, because you will be able to work out for yourself whether a given part of a PDF is something a black rectangle touches. That is a more durable skill than a checklist, and it is the thing that stops the next surprise.

The one idea that explains all of it

A PDF file is a set of numbered objects. Objects are dictionaries of key-and-value pairs, arrays, numbers, strings, or streams of compressed data. They reference each other by number, so the file is really a graph. At the root of the graph sits a catalog, which points to a tree of pages; at the end of the file sits a cross-reference table telling a reader where in the bytes each numbered object lives.

The crucial consequence: the page you see on screen is a rendering, not a thing. When your reader draws page four, it finds page four's object, follows a reference to a content stream, and executes the drawing instructions it contains — move here, select this font, show these glyph codes, paint this image. What you perceive as "the words on the page" are instructions in a stream plus a font object that maps codes to shapes. Nothing in the file is the appearance of the page. The appearance is produced on demand.

Everything that goes wrong with redaction follows from this. Changing the appearance and changing the file are two different operations, and only software that was designed to do the second one does it. A reader that draws an opaque shape over a word has added an instruction; the instruction that shows the word is still in the stream below it, and whoever extracts text from that stream gets the word back. This is the distinction explored at length in permanent redaction versus visual redaction; here it is just the first entry in a longer inventory.

The inventory: what holds what

Below is the practical list of places content lives in a PDF. For each one the question to keep in mind is: does drawing a black box over part of a page change this?

Page content streams. The drawing instructions for a page, including every text-showing operator. This is where body text, table values and captions live. A visual cover-up leaves them intact. A real redaction either edits this stream or replaces it.

Fonts and their encodings. A content stream shows glyph codes, not letters. The font object translates. Embedded font subsets and character-mapping tables are what let a tool turn codes back into readable text — and when those tables are unusual, extraction gets lossy, which is why a text search sometimes fails to find a string that is unmistakably on the page. Absence of a search hit is evidence about the font, not about whether the text is still there.

Image XObjects. Pictures referenced by a page. A scan is one big image per page; a report might embed a chart. An image is opaque to text search but entirely readable to a human and to OCR, so a sensitive value inside a picture is not protected by being a picture.

Form XObjects. Reusable blocks of content drawn by reference, often a letterhead, a stamp, or a repeated table. The same block can be drawn on many pages from one definition, which means a value inside it may appear in more places than you remembered to check.

Annotations. Comments, highlights, sticky notes, stamps, link areas and markup live in a per-page array and are separate from page content. They render on top of the page in most viewers, and some viewers hide them by default. A box painted on the page does nothing to them.

Form field values. An interactive form stores the value of a field as data, separately from the appearance stream that displays it. The two can disagree. Covering what a field looks like leaves the stored value available to anything reading the form data rather than the page.

The document information dictionary and the XMP metadata stream. Title, author, subject, keywords, creator and producer, plus a parallel metadata packet that can carry considerably more. This is document-level and has no relationship to any page, so it is untouched by anything you do to page content.

The outline tree. Bookmarks. Visible, user-facing, structured text in the navigation pane, holding exactly the chapter and exhibit labels people name after the thing they are about.

Embedded files. PDF can carry arbitrary attached files — the source spreadsheet, an email export, the original word-processor document. An attachment is a complete unredacted file sitting inside the one you are redacting.

Optional content groups. Layers that can be switched on and off. Content on a hidden layer is present in the file and visible to any recipient who turns the layer on.

Actions and document-level scripts. Behaviour attached to opening the document, clicking an area, or recalculating a form. Scripts can reference values that never appear on a page.

The structure tree of a tagged PDF. Accessibility tagging stores a logical reading order and alternative text for figures. Alt text is written text that describes an image — including, sometimes, describing it accurately enough to matter.

Prior generations of objects. A PDF can be updated by appending to the end of the file instead of rewriting it, leaving the superseded objects in place while new cross-reference data points past them. Written that way, the file contains its own earlier state. Whether yours does depends on how your software last saved it, and you generally cannot tell by looking.

Several of these non-page surfaces get their own treatment in can sensitive data hide in PDF comments, bookmarks, or attachments. The point of listing them together is the pattern: of roughly a dozen places content can sit, exactly one is the page content stream you are drawing on.

What the tool on this page does to that graph

Worth stating precisely, because it is unusual among the options and it is all verifiable in the client-side code this page loads. The export does not edit your PDF. It creates an empty PDF document, then for each page of the original it renders that page to an offscreen canvas at a fixed internal scale, paints your black rectangles onto the canvas, encodes the result as a JPEG, and adds that JPEG as a full-page image in the new document at the original page's dimensions. One output page per input page. At the end it sets the new document's own title to "Redacted Document" and its producer and creator to "HidePDF".

Read that against the inventory and the consequences fall out mechanically. Nothing from the original object graph is copied, so there is no path by which content streams, fonts, annotations, form field values, bookmarks, embedded files, layers, scripts, structure tags, the original metadata or any earlier generation of the file could appear in the output. They are not stripped item by item — they were never carried over. The covered content is gone for the same structural reason: the rectangle is painted onto the canvas before the image is encoded, so the pixels underneath are not in the data.

The honest flip side is that the output loses things you might want. Selectable text, working hyperlinks and accessibility tags are objects too, and rebuilding from images does not recreate them. The page becomes a picture — legible to a reader and to OCR, but no longer searchable. That trade is the right one when the goal is withholding content and the wrong one when the document needs to stay machine-readable. It all happens in the browser tab, with no upload required, which is what you want while the unredacted original is open.

Reading your own file's anatomy

You do not need to parse PDF syntax to audit a document. A reasonable pass, in order:

  1. Open every pane, not just the page. Comments, attachments, bookmarks, layers and form fields each have their own panel in a desktop reader. Content you have never looked at is content you have never checked.
  2. Open document properties. Title, author, subject and keywords are right there, and they are frequently the name of a person or a matter.
  3. Select all on the page and paste it somewhere. What lands in the clipboard is what the content stream still holds. This is the single fastest test of whether a cover-up is a redaction.
  4. Compare file size against apparent content. A three-page document weighing many megabytes is carrying something — an attachment, a large image, or a long update history.
  5. Judge the finished artefact, never the editor. Close the file, reopen the thing you are about to send, and look at it at high zoom. Every verification that matters happens on the export.

Common mistakes and misconceptions

Treating "I can't find it" as "it isn't there". A failed text search can mean the font's character mapping defeated extraction, or the text is in an image, or the viewer indexed only part of the file. Negative results from one tool are weak evidence.

Believing the page is the document. This is the root error and it wears many disguises: covering a form's displayed value, blacking out a page that has a comment thread, redacting a body paragraph while a bookmark names the subject.

Assuming an image is safe because it is not text. Opacity to search is not opacity to reading. A legible number in a chart is a disclosed number.

Thinking metadata is page-level. The information dictionary and XMP packet belong to the document. No amount of work on page seven changes them.

Confusing flattening with removing. Flattening collapses layers or annotations into page content. Depending on what the software does, that can bake a comment permanently into the picture rather than deleting it — which is the opposite of what the person clicking the button usually wanted.

Assuming encryption substitutes for redaction. A password controls access to a file; it does nothing about what the file contains once it is open. The content is still in the objects.

Forgetting that a rebuilt file loses structure on purpose. If a recipient needs to search or an accessibility requirement applies, a flattened export is a problem to plan for, not a surprise to discover after sending.

Related guides

See also, or use the redaction tool above:

Frequently asked questions

Do I need to understand PDF internals to redact a document safely?

No, but you do need one idea from them: a PDF is a collection of objects, and the page you look at is only some of those objects. Almost every redaction failure is the same mistake in a new costume — someone changed what the page looks like and assumed that changed the file. If you hold on to the distinction between the appearance of a page and the objects that produced it, you will ask the right follow-up question in situations nobody wrote a checklist for. You do not need to be able to read PDF syntax, name the object types, or open the file in a hex editor. The practical version of the knowledge is a habit: after you redact, inspect the artefact you are about to send rather than the screen you redacted on, and open the panes of your reader that show things other than the page — comments, attachments, bookmarks, layers, document properties.

Why does rebuilding a PDF from page images remove so much at once?

Because nothing is copied across. The export on this page creates an empty PDF document, renders each page of your original to a canvas, paints your black rectangles onto that canvas, encodes it as a JPEG and adds it as a full-page image in the new document, one output page per input page at the original page dimensions. The original file's object graph is never touched by that process, so there is no mechanism by which a text layer, an annotation, a form field value, a bookmark title, an embedded attachment or the original metadata could survive into the output — they are not removed one by one, they are simply never carried over. The new document's own metadata is set explicitly in the code to a title of Redacted Document with HidePDF as producer and creator. The same property is why the export also loses things you may have wanted: selectable text, working links and accessibility tags are objects too, and they are not recreated.

Can an older version of the document still be inside the file?

It can, depending entirely on how the file was last written. PDF permits a file to be updated by appending changes to the end rather than rewriting it, which leaves the superseded objects present in the earlier bytes while the newer cross-reference data points at the replacements. A file written that way carries its own prior state, and a reader that follows the older references can reach content the current version does not display. This is a property of the file, not of any particular editing step: a save that fully rewrites the document does not leave that history, and an appending save does. Because you usually cannot tell by looking which kind of save your software performed, treat it as a reason to send a freshly rebuilt file rather than an edited original — a document assembled from scratch has no earlier generation to recover.

If the page is now an image, is anything still readable in the file?

The pixels are. Flattening a page removes the text objects that spelled out a word, but it does not make the remaining picture unreadable — a person reads it, and so does optical character recognition. That matters in two directions. Content you covered is gone from the image because the rectangle was painted before the image was encoded, so there is nothing under it to recover. Content you did not cover is fully legible, including anything you assumed was obscured by being faint, small, low-contrast, or partly behind something else. The practical check is to look at the finished export at high zoom rather than at the editing view, because faintness and overlap are the two things that routinely survive a flatten while feeling redacted at normal size.