Skip to content
Product6 min read

Scanned Documents and OCR in ONYRI Sanitize: How It Works, and Its Limits

ONYRI Sanitize reads scanned PDFs and photos with on-device OCR in English, French and German, then blocks any page it cannot read safely before masking it.

By Alexis de ONYRI
See it on your own document

A scanned PDF or a photo is just pixels, with no text hiding underneath to select or copy. ONYRI Sanitize reads that image using on-device OCR, short for optical character recognition, then masks whatever it finds. The whole read happens inside your browser. If a page turns out too blurry to trust, ONYRI blocks the download instead of handing you a file that still shows the name.

How Does On-Device OCR Work?

OCR looks at the shapes on a page and turns them into real letters and numbers. ONYRI runs this step in your browser, on the open-source engine Tesseract, the same technology many privacy-minded tools use. Nothing you scan gets uploaded anywhere. The image, the recognized text and the masked result all stay on your device, and the mapping disappears the moment you reload the page.

  • You add a PDF, a Word file, or a photo.
  • ONYRI checks each page. A page with a real text layer is read directly, no OCR needed.
  • A page with no text layer, like a scan or a photo, goes through OCR instead.
  • OCR reads English, French and German at the same time, automatically.
  • The recognized text feeds the exact same detectors used on a typed document.

Which Pages Actually Get OCR?

Not every page needs OCR. A text PDF, one where you could already select and copy a sentence, has a text layer ONYRI reads directly. That is faster and more accurate than OCR. OCR only kicks in for a PDF page with no usable text layer. It also runs for any image you add, such as a phone photo of a lease or a scanned ID card.

What Can This OCR Not Read?

OCR has real limits, and ONYRI does not hide them. Handwriting is not read at all, so a handwritten note or a signed line stays exactly as scanned. Scan quality matters just as much. A crooked photo, a dark room or a blurry lease slows OCR down and makes it miss more. Right now, the engine supports English, French and German, and nothing else.

SourceWhat it recommends for a text page
Tesseract OCR project documentationAt least 300 dots per inch, with the page kept as straight as possible
The National Archives (UK), digitisation guidance300 pixels per inch by default for an ordinary document
A typical phone camera photoOften well above 300 dpi already, but blur and a crooked angle hurt OCR more than resolution does

How Do You Scan So More Gets Detected?

  1. 1Lay the page flat. A page curved inside a bound notebook blurs the lines OCR needs.
  2. 2Keep it straight. The National Archives asks for a skew of about one degree at most. Tesseract's own documentation says a skewed page badly hurts OCR quality.
  3. 3Use good light and real contrast. Dark ink on plain white paper reads far better than a faded photocopy of a photocopy.
  4. 4Scan at 300 dots per inch or more for text. A lower setting saves space but throws away the detail OCR relies on.
  5. 5Fit the whole page in the frame. A cropped edge can cut off the exact data you meant to mask.

What Should You Check Before You Download?

OCR is not the last word. ONYRI shows a page-by-page preview and a list of every value it detected, grouped by family, such as names, emails or ID numbers. Each value has its own checkbox, so Jane Example can leave her own name visible on a cover page if she wants to.

  1. 1Read every page of the preview, not just the first one.
  2. 2Open the detected data list and scan it for a false match or a missed value.
  3. 3Watch for the warning that a value could not be placed. Check that spot with your own eyes.
  4. 4Switch between Black marker and Token mode if the first choice does not fit.
  5. 5Download once the preview looks right. The export is flattened, so there is no editing it afterward.

OCR gets a scan closer to a normal document, but masking still is not perfect. Detection is not exhaustive, a rough scan can still hide a name from the engine, and reading the preview stays your job. Treat OCR as a head start, never as the last check.

Frequently asked questions

Does ONYRI Sanitize upload my scanned document to a server?
No. OCR, like the rest of the pipeline, runs inside your browser. The server only receives usage counters, such as the number of documents or pages processed, never the file, its text or a detected value.
Can the OCR read handwriting?
No, it reads printed and typed text only. A handwritten note, a signature or a hand-filled form field stays exactly as scanned, so check those spots yourself before you share the file.
Why did my document get blocked instead of anonymized?
A page blocks the export when it has no usable text layer and OCR cannot make sense of it either. This is a safety rule. ONYRI would rather stop than hand you a file that looks masked but still shows the original text underneath.
Which languages does the OCR support?
English, French and German, all at once. ONYRI checks all three every time you scan, so there is nothing to choose beforehand.
Does a higher-resolution scan always detect more?
Mostly, up to a point. Groups such as Tesseract and The National Archives put the useful minimum around 300 dots per inch for text. Past that, a straight, well-lit, sharp scan matters more than piling on extra resolution.

Sources & references

On this siteRedact a scanned document or an image

Mask a document without uploading it

ONYRI Sanitize finds names, identifiers, bank details and secrets in a PDF, a Word file or a scan, and masks them in your browser. You check the preview, then download a flattened copy.

Read next