Consistent Pseudonyms: How to Replace Names Without Losing Meaning
Role labels, numbered tokens, or fake names: three ways to replace real names in a document while keeping one consistent label per person throughout.
A black bar over a name breaks a story. If a report says the manager ignored three complaints, the reader needs to know it is the same manager every time, not three different people. Consistent pseudonyms fix this: they hide who someone is, and they keep the text readable and connected.
Why Black Bars Break Meaning
A black bar treats every name the same way: it disappears. But a training case, a research transcript, or an AI prompt often needs the reader to follow one person across many pages. If Jane Example becomes an unreadable bar three times, then a different bar the fourth time, the story falls apart. The reader cannot tell if it was one manager or four.
Three Ways to Replace a Name
| Method | How it works | Good for | Watch out for |
|---|---|---|---|
| Role label | Replace the name with a function, like the tenant or the manager | Short texts, few people, plain-language summaries | Two people with the same role become confusing |
| Numbered token | Replace each name with a code, like NAME1, repeated for that person | Long documents, many people, training data, AI prompts | The token list itself must be protected |
| Realistic fake name | Replace the real name with another invented name, like Jane Example | Readable case studies, workshops, sample contracts | Readers may confuse it with a real living person |
Pick one method per document and stick to it. Mixing role labels and tokens in the same text usually confuses more than it protects. The reader has to learn two different codes at once.
The One Rule That Keeps a Text Coherent
The rule is simple: one person, one label, on every page. If Jane Example becomes [NAME2] on page one, she stays [NAME2] on page ten. Break this rule once and the reader loses the thread, or worse, assumes two different people said the same thing.
- Build the label list before you start reading the document, not while you go.
- Keep the same label for a person even if their name is spelled two different ways in the original text.
- Drop gender and role from a label unless the text needs it to make sense. Keep the manager, for example, if the case is about a management decision.
- Reusing a label across two unrelated documents can link them together, so treat each project's label list separately.
Generalizing Details Without Losing Context
A label hides a name, but the sentence around it can still identify someone. The 34-year-old sales director in Annecy narrows the search to one person, even without a name. Generalization keeps the sentence useful, and it widens the group the sentence could describe.
- Turn an exact age into a band: 34 years old becomes in her thirties.
- Turn a small town into a region: Annecy becomes a town in the French Alps.
- Turn an exact date into a shifted or rounded one. 14 March 2024 becomes spring 2024, shifted the same way for that person across the whole text.
- Turn an overly specific job title into a broader one when the title alone identifies someone. The only VP of Engineering becomes a senior engineering manager.
Is a Pseudonymized Text Still Personal Data
Under EU law, yes. Article 4(5) of the GDPR defines pseudonymisation as processing that keeps data from being linked to a person without extra information. That extra information must be kept separately and protected. Recital 26 adds that pseudonymised data can still identify someone once that extra information comes back. So it still counts as personal data about an identifiable person.
The European Data Protection Board adopted draft Guidelines 01/2025 on Pseudonymisation in January 2025. They describe pseudonymisation as a safeguard, not a way to escape GDPR obligations.
If building this label list by hand feels slow, ONYRI Sanitize's Token mode (Pro plan) does the consistency part for you. Each detected name, email address, or phone number gets one label, such as [NAME1], repeated everywhere it appears in the document. The tool keeps no key file to restore the original values, so save your unmasked original before you share the labeled version.
A Quick Checklist Before You Publish
- Pick one method, role label, numbered token, or fake name, for the whole document.
- Give every person one label and check it stays the same from the first page to the last.
- Generalize any detail, age, place, date, or title, that could single someone out on its own.
- Store the label list separately from the pseudonymized text, or delete it if you will never need to reverse it.
- Remember pseudonymised text is still personal data under GDPR, so it still needs a legal basis and reasonable security.
Frequently asked questions
- What is the difference between pseudonymization and anonymization?
- Pseudonymization replaces identifying details with labels but keeps a way back to the original, even if that way is stored separately. Anonymization removes that way back entirely. Under EU law, pseudonymised data is still personal data. Properly anonymised data is not, according to GDPR Recital 26.
- Can I reuse the same fake name, like Jane Example, across several unrelated documents?
- It is safer not to. If the same fake name always maps to the same real person across projects, it creates a link between them. Someone who sees two documents could then narrow down who Jane Example really is. Use a separate label list for each project.
- How specific can a role label be before it identifies someone?
- A role label is only safe if more than one person could hold it. The regional director is fine in a large company. In a three-person team, it points to one name. When a role has one obvious holder, use a numbered token or a fake name instead.
- Do I need to keep the file linking labels to real names?
- Only if you will genuinely need to reverse the pseudonymization later. If you keep it, store it apart from the pseudonymized text, on a different system, and limit who can open it. If you never need to reverse it, deleting it removes one more thing that could leak.
Sources & references
- Guidelines 01/2025 on Pseudonymisation (draft) — European Data Protection Board (EDPB)
- Regulation (EU) 2016/679 (GDPR), Article 4(5) and Recital 26 — EUR-Lex, Publications Office of the European Union
- Pseudonymisation Techniques and Best Practices — European Union Agency for Cybersecurity (ENISA)
Mask a document without uploading it
ONYRI Sanitize finds names, identifiers, bank details and secrets in a PDF, a Word file or a scan, and masks them in your browser. You check the preview, then download a flattened copy.