Skip to content
Product7 min read

How accurate is ONYRI Sanitize? Our detection benchmark on 48 fictitious documents

We tested our detection on 48 fictitious documents: 85.5% of sensitive items fully masked, 93.8% of masks needed. Method, misses and corpus published.

By Alexis de ONYRI

On 48 fictitious documents, ONYRI Sanitize fully covered 85.5% of the 869 items a careful person would mask. And 93.8% of its 872 masks fell on something that needed masking. Both figures measure the text of documents we wrote ourselves, with the detection rules only.

The first figure is the recall: the share of what had to be hidden that was hidden. The second is the precision: the share of masks that were needed. NIST (SP 800-188, September 2023) says agencies can adopt a standard with measurable performance levels when they de-identify government datasets. The idea holds here too: “accurate” needs a number and a method.

What is in the test corpus?

The corpus holds 48 fictitious documents, about 59,000 characters in all, with 26 to 62 lines each. They are written in six languages: English, French, German, Spanish, Italian and Dutch. They come from 11 countries. Luxembourg has one French and one German document.

The kinds include payslips, contracts, invoices, bank statements, medical letters, insurance claims, leases, CVs, e-mail threads, minutes and court decisions. A few imitate a scan read by OCR, with stray spaces and the letter O for a zero.

We wrote them for this purpose, with the help of an AI writing assistant. The first 36 were written without looking at the detectors' code. The 12 Spanish, Italian, Dutch and Luxembourg documents came with those countries' detectors. Then we annotated the sensitive items, following written rules. People and companies are invented. E-mail addresses and links use domains reserved for examples. Numbers follow each country's real format, with valid check digits (digits computed from the others, to catch typos) where the format has them.

The corpus also holds traps: things that look sensitive but must stay visible. A mask on a trap counts against us. In all, 869 items are annotated, in 17 categories. The traps include:

  • Streets named after people
  • Cities that are also first names
  • Public figures
  • Capitalised common words
  • Product codes
  • Article numbers of laws
  • Percentages
  • Years on their own

How do we score it?

  • Fully masked (recall): the share of annotated items that were hidden. An item counts only if every letter and digit of it is under a mask. A half-masked IBAN is missed.
  • Right label: the share of items touched by a mask of the matching category, such as a phone mask on a phone number. For the first figure the label does not matter. A mask is a mask.
  • Precision: the share of masks that touch at least one annotated item. A mask that touches none is a false positive, and it counts against us.

Recall and precision follow the definitions in Google's Machine Learning Crash Course. We measure only the text: the engine's detection rules.

What does it find, by type of data?

Six categories reach 100%, but four of them rest on only 4 to 7 items.

CategoryItemsFully maskedRight label
Person names15085.3%94.7%
E-mail addresses56100%100%
Phone numbers5396.2%96.2%
Postal addresses7680.3%97.4%
Dates17188.3%87.1%
National ID numbers2387.0%91.3%
Bank details3889.5%76.3%
Card numbers6100%100%
Amounts with a currency137100%100%
Company IDs2090.0%75.0%
Reference numbers4156.1%53.7%
Organisation names6464.1%64.1%
Medical codes (ICD-10)7100%100%
Medical terms in words150%0%
Secrets (passwords)450.0%50.0%
Private URLs4100%100%
IP addresses4100%100%
Detection by category (run of 6 October 2026, all detectors, country known)

Two weak spots stand out: reference numbers (56.1%) and organisation names (64.1%). The right label trails the fully masked share for bank details (76.3% against 89.5%) and company IDs (75.0% against 90.0%). The value was hidden, but under another label. In Token mode the label is part of the replacement, so a wrong one shows.

How do the results differ by country and language?

CountryDocumentsItemsFully maskedPrecision
France815388.2%98.6%
Belgium44582.2%93.5%
United States713577.8%81.6%
United Kingdom59171.4%95.2%
Germany612688.9%94.5%
Austria34877.1%97.6%
Switzerland35096.0%98.0%
Spain47483.8%96.1%
Italy35896.6%93.8%
Netherlands35394.3%96.6%
Luxembourg236100%95.5%
Detection by country, all detectors

By language, French (88.0%) and German (88.8%) lead the larger groups on recall. Spanish, Italian and Dutch score 83.8% to 96.6% recall, on only 3 or 4 documents each. English is the weakest: 75.2% recall and 86.8% precision over 12 documents.

The United States has the lowest precision (81.6%), partly because many first names there are also words or places. The United Kingdom has the lowest recall (71.4%), partly because no detector covers UK street addresses or sort codes (bank branch numbers). Seven countries have only 2 to 4 documents each, so one item moves their figures by one to three points.

What does it miss, and why?

It missed 126 items: organisation names 23, person names 22, dates 20, reference numbers 18, medical terms 15, addresses 15 and 13 more. The main causes:

  • Medical terms in words (15 items, 0%). No detector exists for a diagnosis or a treatment in plain text.
  • Organisation names with no legal form (64.1%). The rules catch a name followed by GmbH, Ltd or SA. A bakery or a clinic with an ordinary name and no form slips through.
  • Reference numbers with a label we do not list (56.1%). A number is found when a listed keyword, such as “Our ref” or “policy no.”, sits right before it. “Ticket” is not listed.
  • Rare names. A first name or surname missing from our name lists can leave half a name visible.
  • UK addresses. UK street addresses and sort codes have no detector.
  • Incomplete dates. Years alone in one CV (we annotated them in this CV only), and a day and month with no year in a bank statement.
  • Scan noise. In the simulated scans, an O for a zero or stray spaces broke some dates, names and addresses.
  • A password inside a sentence, with no “password:” right before it.

Every gap can be closed by hand. Use “Select text”, “Draw an area” or “Search the document”, then mask all results, or type a missed term under “Also mask”.

What are the false positives?

Among its 872 masks, 54 fell on nothing we had annotated. Of the 54, 38 sit on person names. The main causes:

  • First names that are also words or places (Bill, Dallas, Charlotte, Florence).
  • Names and dates we annotated elsewhere in the document but not here, such as “Ms. Moreno's” (14 of the 54). These masks are right. Our annotation missed them.
  • Public figures (a street named after Victor Hugo).
  • Generic mailboxes such as payroll@, left unannotated in three documents but annotated in others.
  • All-caps words that look like a bank code.
  • Words damaged by scan noise and taken for secrets.

Should a redaction tool accept some over-masking? We think so. A false positive hides a harmless word. A miss leaks a value. Google's course advises favouring recall when false negatives cost more than false positives. Over-masking still has a cost: too many useless masks tire the reviewer. So you can uncheck any value in the “Detected data” list, and a test guards precision.

What does the Default profile score?

On the Free plan, you get the Default profile. It leaves out the technical-secret detectors, which are Pro, and it guesses the country from the text. It scores 85.2% recall and 94.2% precision, with 50 false positives out of 865 masks. Reference numbers fall to 53.7%, and secrets to 0%.

What do these figures not say?

The first version of the corpus, from 27 September 2026, had 36 documents in English, French and German. Its first run found 73.6% with 91.7% precision. That day we added 12 documents from Spain, Italy, the Netherlands and Luxembourg.

How do we use the benchmark?

The benchmark is part of the test suite we run before each release. The test fails if any category loses recall against a stored reference, or if precision drops by more than one point.

Can you check our numbers?

Yes. We publish the 48 documents with their inline annotations, the annotation and scoring rules and the results. The licence is Creative Commons Attribution 4.0 International (CC BY 4.0). Use them to test any tool, ours included.

Frequently asked questions

Does 85.5% recall mean I can skip the review?

No. About one item in seven was not fully masked on this corpus. Always read the preview, use the “Detected data” list, and mask what is left by hand.

Is my document sent anywhere when I use the app?

No. Detection and masking run in your browser, and the file is never uploaded. The server receives only counters: the document kind, the number of pages and the number of masks.

Is the name-recognition model part of these figures?

No. The benchmark measures the detection rules only. On a computer with a mouse or trackpad, a name-recognition model that runs on your device usually also looks for person and company names. It complements the rules and is not exhaustive.

Does masking at this level make a document anonymous?

No. Hiding what a tool detects reduces exposure. The UK ICO's anonymisation guidance (28 March 2025) says you should reduce the risk of identifying people to a sufficiently remote level. Token mode is pseudonymisation. Anyone who holds your original can link the labels back, and the ICO says pseudonymised data is personal data in such hands.

Sources & references

  1. Classification: Accuracy, recall, precision, and related metricsGoogle for Developers, Machine Learning Crash Course
  2. NIST SP 800-188: De-Identifying Government Datasets: Techniques and Governance (September 2023)NIST
  3. Anonymisation guidance (28 March 2025)Information Commissioner's Office (ICO)
  4. PseudonymisationInformation Commissioner's Office (ICO)

Mask a document without uploading it

ONYRI Sanitize finds names, identifiers, bank details and secrets in a PDF, a Word file or a scan, and masks them in your browser. You check the preview, then download a flattened copy.