How to Redact a PDF Properly (and Permanently)
Published by PDFToolz · Updated · Technical references listed below
To redact a PDF, you delete sensitive content from the file so nobody can recover it. Covering it with a black box isn't enough. Done wrong, the words look gone on screen but the original text is still underneath, and anyone can copy and paste it.
This guide explains what redaction does, why black boxes and highlights fail, what a proper redaction must remove (visible text, images, the hidden OCR layer and metadata) and how to check the file before you send it.
What PDF redaction means
Redaction means deleting content so that no method can bring it back. In a PDF, sensitive data can sit in several places at once: the visible text and graphics, a hidden text layer on scanned pages, embedded images and metadata inside the file.
The test is simple. After redaction, the words, numbers or images you removed must not exist anywhere in the file. Opening it in a text editor, searching, copying the page or reading the raw bytes must all come up empty. Anything less is hiding, not redaction, and hiding fails.
- Redaction deletes the content from the file for good.
- Concealment hides it behind a shape or color, but it is still there.
- Only real redaction is safe to publish.
Why black boxes don't redact a PDF
The most common mistake is treating redaction as a drawing task. A black rectangle, a black highlight, text colored to match the background or an image pasted on top all look convincing. The original text is still there. Select the page, copy it, and the hidden words paste into any text editor.
This happens in real cases. In January 2019, lawyers for Paul Manafort filed a US federal court document whose blacked-out text could be read by copying and pasting it. In December 2009, the US Transportation Security Administration posted its airport screening manual online with redactions that could be removed to show sensitive procedures. In both cases a box was drawn over the text instead of the text being deleted.
- A black box or shape over text. The text stays in the content stream.
- A black or opaque highlight. It is an overlay, not a deletion.
- Font color set to blend into the page.
- An opaque image placed over the words.
- Any method where the page looks right but nothing was deleted.
What a real redaction removes
Proper redaction deletes content. It doesn't paint over it. A real redaction tool removes the characters, vector shapes and image pixels in the marked area, then draws a black bar over the empty space. The PDF standard (ISO 32000) defines a redaction annotation, but it only marks content for removal. The data stays in the file until you apply the redaction as a separate step.
The visible page is only part of the job. The same information can sit in metadata, comments, form field values, bookmarks, scripts, attachments, hidden layers, image EXIF data and the invisible OCR text layer on scanned pages. The OCR layer is the one people miss most. Any of these can leak what you thought you removed.
- Visible text and vector graphics in the content stream.
- Embedded images, including their EXIF and GPS data.
- The invisible OCR text layer behind scanned pages.
- Document properties and XMP metadata, such as author, title and keywords.
- Comments, annotations and saved form field values.
- Hidden layers (optional content), attachments and bookmarks.
How to redact a PDF step by step
Use software with a real redaction feature that marks content and then deletes it. Editors that only let you draw shapes won't do. The workflow is the same in any proper tool. Mark what to remove, then apply the marks to delete it.
For scanned pages, or when you need full certainty, turn the redacted pages into plain images. This removes any leftover text layer, but the text on those pages can no longer be selected. PDFToolz Redact PDF works this way. It redraws each marked page as an image with the black boxes burned in, so the text and images underneath are gone.
- Work on a copy. Never redact your only original.
- Mark every instance, and search the document to catch repeats.
- Check headers, footers, captions and repeated data.
- Apply the redaction so marked content is deleted, not covered.
- Run a sanitize or 'remove hidden information' step to strip metadata, comments, layers and attachments.
- Save as a new file, then close and reopen it before checking.
Key takeaways
- ✓Redaction means deleting sensitive content for good, not hiding it behind a black box.
- ✓Shapes, highlights and recolored text leave the original text recoverable by copy and paste or search.
- ✓A real redaction removes visible text, images, the invisible OCR layer and hidden metadata.
- ✓Use a tool that deletes the marked content, then run a separate sanitize step for hidden data.
- ✓Before you share, copy the text, search for redacted terms and inspect the metadata.
Tools for the job
Frequently asked questions
Why can I still copy text from a redacted PDF?
Because the redaction was only an overlay, such as a black box, highlight or recolored text. The original characters are still in the content stream, so copying the area returns them. Only real redaction, which deletes those characters, prevents this.
Is a black box over text enough to redact a PDF?
No. The box is a separate object drawn on top, and the text underneath stays in the file. Anyone can copy the covered text, delete the box in an editor or read it in the raw file. Use a tool that removes the content under the mark.
Does redacting a PDF remove its metadata?
Not on its own. Applying redaction to the page leaves document properties, XMP metadata, comments and attachments untouched, and any of them can hold sensitive data. Run a separate sanitize or 'remove hidden information' step, then save to a new file.
How do I redact a scanned PDF?
Scanned PDFs often have an invisible OCR text layer behind the image. Covering the image isn't enough, because that layer stays searchable and copyable. Use a tool that removes both the image area and the OCR text, or turn the page into a flat image after redacting. PDFToolz Redact PDF does the second.
How can I check a PDF was redacted correctly?
Select all text and paste it into a plain text editor, then search the document for each redacted term. Both should return nothing. Zoom into each black mark to check for show-through, and inspect the metadata panel before you share the file.
Can redacted content be recovered from a PDF?
Not if it was truly redacted, meaning deleted from the content stream and the metadata. If it was only covered with a shape, highlight or image, yes. Anyone can recover it by copying, editing or inspecting the file.
Related terms
Sources and further reading
- Redact and sanitize PDFs from Adobe Acrobat Help
- Sanitize PDFs and remove hidden content from Adobe Acrobat Help
Browse all PDF guides, look up a term in the PDF glossary, or head back to the PDFToolz toolkit.