Use a tool that rewrites the file itself rather than one that draws over it. Black rectangles added in a PDF viewer are drawn on top of the text, and the text underneath is still selectable, still copyable, and still read verbatim by any AI you upload it to — this is how most published redaction failures have happened. Proper redaction opens the document, finds the personal data in the actual content, replaces it with tokens, and writes the file back with its layout, styles, tables and formulas intact, so the result is still a usable document. Scanned pages and photographs need one step before that: there is no text in the file at all, only pixels, so the page has to be read with OCR before anything can be found or masked.
The failure everyone repeats
A black box in a PDF is a drawing. It has no relationship to the words beneath it. Select the region, copy, paste — the names come out. Every model you upload that file to reads the text layer directly and never sees your rectangle at all. The same is true of highlighting in white, of covering a cell with a shape in Excel, and of setting text to white in Word.
The test takes five seconds: open the redacted file, select all, and paste into a plain text editor. If you can read a name, so can the model.
What "format-preserving" has to mean
Getting the data out is easy if you are allowed to destroy the document. The work is in giving back something people can still use.
- Word: styles, headings, tables, headers and footers, tracked changes and comments — comments and document properties are where reviewer names hide.
- Excel: formulas that still point at the right cells, sheet names, and the defined names and hidden columns nobody thinks to look at.
- PowerPoint: speaker notes, which are pure text and frequently full of names.
- PDF: text genuinely replaced in the content stream, not covered; and the metadata, which carries the author and often the original filename.
Scans and photographs
A scanned contract is an image. There is no text to find, so no text tool will find anything, and a redactor that reports "0 items masked" on a page full of personal data is worse than one that refuses, because you will believe it. These need OCR first: read the page into text, locate the values, then mask the regions of the image itself.
Tokens, not black boxes
There is a real choice about what to put in place of a name. A black box destroys the structure of the sentence, and a model handed a document full of black boxes produces noticeably worse work. A token keeps the grammar intact and keeps identity consistent — [PERSON_1] in paragraph two is the same human as [PERSON_1] in the appendix, which is exactly the relationship a summary needs to get right.
Check what came back
Whatever tool you use, look at the output before you upload it. The failure mode that matters is the confident one: a clean-looking report, a plausible count, and a name still sitting in the speaker notes. A tool that offers a way to say "you missed this" — and folds that back into what it looks for next time — is one that expects to be wrong occasionally, which is the correct posture.
Common questions
Does drawing a black box over text in a PDF remove it?
No. The box is drawn on top; the text underneath remains selectable and copyable, and any AI you upload the file to reads it verbatim. This is the cause of most published redaction failures. Test it by opening the file, selecting all, and pasting into a plain text editor — if you can read a name, so can the model.
Can you redact a scanned PDF or a photograph?
Only with OCR first. A scan contains no text, just pixels, so a text-based tool finds nothing and may report zero items masked on a page full of personal data. The page has to be read into text, the values located, and then the regions of the image itself masked.
Is it better to replace names with black boxes or with tokens?
Tokens, if the document is going to an AI. A black box destroys the sentence structure and the model produces worse work from it. A token keeps the grammar intact and keeps identity consistent across the whole document, so a summary can correctly say that the person in paragraph two is the person in the appendix.
Try it on your own data
A free workspace takes one step — 250 requests a month, three connected systems, no card.
Create a workspace