How to test a redaction tool before you trust it
A one-afternoon protocol to test any redaction tool: fictitious documents, an answer key, a results table and a check of the exported file.
Test a redaction tool on fictitious documents that look like yours, with every answer marked before you start. Count what it missed first, because a miss leaks, then what it hid by mistake. Finally, open the exported file itself, not the screen, and see what it still contains.
How do you build a test set?
Take the kinds of documents you really handle: contracts, payslips, letters, scans, forms. Write 10 to 20 fictitious versions with invented people, such as Jane Example and Claire Exemple. Do not start from a real file and swap the names, because the other details stay real.
The set should include:
- Your own country's formats for ID numbers, IBANs, phones and addresses.
- A scan, and a PDF with filled form fields.
- A Word file with comments and tracked changes.
- Headers, footers and tables where names appear.
- Traps that must stay visible: a street named after a person, a city that is also a first name, a public figure, a product code.
Use numbers with valid check digits. Every IBAN, most card numbers and many ID numbers carry check digits calculated from the other digits. A tool may skip a number that fails the check. Random digits would give a false picture.
How do you write the answer key?
Write the answer key before you run the tool, never after. List every item a careful person would mask, document by document, with its category. Written afterwards, the key excuses the tool's misses without you noticing.

Settle the grey cases first, and write each rule down. Typical questions:
- Is a shared mailbox such as payroll@ personal?
- Are all dates personal, or only birth dates?
- Does a job title identify someone in a team of five?
- Does the name of a public figure stay visible?
How do you count the results?
Run every document with the settings you would really use, and save the exported files. An item counts as found only if it is fully covered. A half-masked IBAN still leaks, so it is “partly masked”.
| Category | Items to mask | Fully masked | Partly masked | Missed | Masked by mistake |
|---|---|---|---|---|---|
| Names | 40 | 36 | 2 | 2 | 3 |
| Addresses | 40 | 36 | 2 | 2 | 1 |
| Bank details | 10 | 5 | 3 | 2 | 0 |
| ID numbers | 10 | 9 | 0 | 1 | 0 |
| Total | 100 | 86 | 7 | 7 | 4 |
Here the tool fully masked 86 of 100 items. That is a recall of 86%: the share of what had to be hidden that was hidden. It drew 97 masks: 93 needed (86 complete, 7 partial) and 4 by mistake. That is a precision of about 96%: the share of masks that were needed.
How do you read an accuracy claim?
Vendors quote recall, precision, or a blend of both. Google's Machine Learning Crash Course says to favour recall when a miss costs more than a false alarm. In redaction, it does. A miss leaks a name. A false alarm hides a harmless word.
An average can also hide a weak category. Look back at the table: 86% overall, but only half of the bank details were fully masked. Ask for results per category, with the number of misses.
| Question to ask | Why it matters |
|---|---|
| Which documents and languages, scans included? | Clean text says little about your scans, forms and countries. |
| Who wrote the answer key? | A vendor can score well on a test it wrote and tuned itself. |
| Is the test set published? | Only then can you check the claim yourself. |
| Is a partly masked item a hit? | A half-covered number still leaks. |
ONYRI Sanitize publishes its own test set: 48 fictitious documents in 6 languages from 11 countries. It comes with the answer key, the scoring rules and the results by category, under a CC BY 4.0 licence. You can run the same files through any tool, ours included. The limit: we wrote this set and tuned our detectors on it, so build your own too.
How do you check the exported file, not the screen?
A preview shows pixels, and the file can hold more. Lower Saxony's data protection commissioner tells users to check the final document carefully. The UK Information Commissioner's Office (ICO) warns that a recipient may reveal redacted text by pasting a PDF with black rectangles into Notepad. Its anonymisation guidance describes a “motivated intruder”: a reasonably competent person who wants to identify people. Think like that person.
- Select all, copy, paste into a plain text editor, and search for each masked value.
- Search the file for a surname alone, then for the last four digits of a number.
- Open the properties, comments, tracked changes and attachments.
- Read the file name. It may carry a name.
Scans need one more look. The OCRmyPDF documentation says the tool adds text “layers” to images in PDFs, making scanned PDFs searchable. If an exported scan keeps such a layer, masked words may still be in it. Checking one file before sending has its own guide on this site.
Which hard cases should you add?
A tool that copes with clean text can still fail on real life. Add the cases your documents contain.
- A crooked or grainy scan, and a phone photo.
- Handwriting, such as a signed note in the margin.
- A name inside a logo or a signature block.
- A table, where values sit in cells.
- A number split across two lines, such as a wrapped IBAN.
- Hyphenated, accented names, such as Anne-Sophie Lefèvre.
- Headings in capitals, such as JANE EXAMPLE.
How do you decide, and what should you keep?
Set pass criteria per category before you look at the results. For example: every ID number and bank detail fully masked, and a limit on over-masks elsewhere. The US institute NIST, in its 2023 guide on de-identifying government datasets, says agencies can adopt a standard with measurable performance levels. The idea carries over to documents.
No score replaces a person. The ICO advises reviewing redactions, or a sample of them, for example by a peer. Keep that check whatever the tool scored. Then keep a file. It holds:
- The tool, its version and the date.
- The test set, the answer key and the results table.
- The pass criteria, the decision and who took it.
That file is your evidence. The GDPR requires a controller to be able to demonstrate compliance (Article 5(2)) and to review its measures where necessary (Article 24). So run the test set again after each update.
Frequently asked questions
How many test documents do I need?
Start with 10 to 20 of your real kinds. That is enough to see patterns, and small enough to finish in an afternoon. Add more for each new kind of file or country.
Can I test with real documents if I hide the names first?
No. Hiding the names is the very thing you are testing, and the rest of the file is still real. Use fictitious documents.
Does the same protocol work for a chatbot asked to redact a file?
Yes. Judge the output file, not the promise. A hosted chatbot also receives whatever file you give it, so use fictitious files.
What score should I demand?
There is no universal figure. Set a target per category, based on the harm of a leak. Expect to review the output by hand anyway.
Sources & references
- Article 28 GDPR: Processorgdpr-info.eu (Intersoft Consulting)
- Article 5 GDPR: Principles relating to processing of personal datagdpr-info.eu (Intersoft Consulting)
- Article 24 GDPR: Responsibility of the controllergdpr-info.eu (Intersoft Consulting)
- Guidelines 07/2020 on the concepts of controller and processor in the GDPR (version 2.1, 7 July 2021)European Data Protection Board (EDPB)
- Tester vos applications (27 January 2020, in French)CNIL
- How do we avoid an accidental breach when redacting information? (31 July 2025)Information Commissioner's Office (ICO), UK
- How do we ensure anonymisation is effective? (motivated intruder test)Information Commissioner's Office (ICO), UK
- Classification: Accuracy, recall, precision, and related metrics (Machine Learning Crash Course)Google for Developers
- NIST SP 800-188: De-Identifying Government Datasets: Techniques and Governance (September 2023)National Institute of Standards and Technology (NIST)
- Hinweise zum Schwärzen von Dokumenten (July 2025, in German)Der Landesbeauftragte für den Datenschutz Niedersachsen
- OCRmyPDF introductionOCRmyPDF documentation