Tokenization, Hashing or Encryption: Which One Pseudonymizes?
Tokenization, hashing and encryption all get called pseudonymization. Here is what each one does to a value in a document, and which one stays readable.
Tokenization, hashing and encryption all get called pseudonymization, and the words get mixed up. Only some of them let you swap a name or a number for a label a reader can still follow. The rest scramble a value into something unreadable, useful for matching records, not for a document a person reads.
What does pseudonymisation mean under GDPR?
GDPR defines pseudonymisation in Article 4(5). It means processing personal data so nobody can link it back to one person without extra information. That extra information must be stored apart and protected. A pseudonym alone does not remove personal data's legal status.
The European Data Protection Board calls this protected circle the pseudonymisation domain in its 2025 guidelines on the subject. It is whoever should not be able to attribute the data on their own. The technique you pick decides how hard that circle is to break into.
What are the four ways to replace a value in a document?
ENISA, the EU's cybersecurity agency, groups pseudonymisation into a handful of building blocks in its 2019 guide. Four of them come up again and again in tools that mask a document.
The simplest one swaps a value for a label from a counter or a random generator, like turning Jane Example into [NAME1]. Nothing is calculated from her name. A table kept by whoever ran the swap says which label means which value, and that table is the only way back.
A cryptographic hash function turns Jane's email into a fixed string of letters and numbers, always the same string for the same input. Nobody needs a secret to compute it. That is also its problem, as the next section explains.
A keyed hash, often called an HMAC, adds a secret key to the calculation. Without that key, nobody can reproduce the string or check a guess against it. ENISA calls this a robust pseudonymisation technique, as long as the key stays secret.
Encryption locks the original value with a key and can hand it back unchanged through decryption. It is the only one of the four built to be reversed on purpose. Think of a locked drawer, built to be opened again by whoever holds the key.
| Technique | Who can reverse it | Reads like text | Fits a shared document |
|---|---|---|---|
| Counter or random label | Whoever holds the mapping table | Yes | Yes, this is how document tools do it |
| Plain hash | Anyone who can list or guess candidate values | No | No, and risky for names or emails |
| Keyed hash (HMAC) | Only whoever holds the secret key | No | Rare, better for matching rows in a database |
| Encryption | Only whoever holds the decryption key | No | No, the encrypted value does not read as text |
Only the first row keeps a document readable. The other three turn a value into noise, which is exactly what a database needs and exactly what a document does not.
Why is a plain hash of a name or an email weak?
A name or an email address does not come from an unlimited set of possibilities. Someone who wants to break the hash can hash every candidate they can think of and compare the results. They do not even need to guess the hash function itself.
The European Data Protection Board gives this exact example in its 2025 guidelines. Say an attacker knows a dataset was pseudonymised with a plain hash of a name. They can hash every name they already hold elsewhere and see which ones match.
- A secret key turns a plain hash into a keyed hash, far harder to attack without that key.
- A password-hashing function built to run slowly, such as argon2, blocks fast guessing. ENISA and Germany's BSI both recommend this kind of function, per the EDPB guidelines.
- A long, random secret, not a date or a short word, gives an attacker nothing worth guessing.
What does tokenization mean in the payment-card sense?
In payment security, tokenization has a narrower meaning. The PCI Security Standards Council sets the rules for card payments. It defines a token as a surrogate value that replaces a card number, and de-tokenization as the step that gets the card number back.
A token does not have to come from hashing the card number. The Council's guidelines list an index, a sequence number or a random number as common ways to generate one too. Each token is matched to the original in a secure table called a card data vault.
That makes payment tokenization closer to the counter and random-label family than to hashing or encryption. Its safety comes from keeping that table locked up, not from any secret maths behind the token itself.
Which technique fits a document you share for reading?
A document gets read by a person, so its masked values need to stay short and consistent. [NAME1] has to mean the same person every time it appears, and a reader still has to be able to follow the sentence around it.
This is why ONYRI Sanitize's Token mode, on Pro plans, builds labels from a counter rather than a hash. Each value becomes a short tag such as [NAME1] or [MAIL1], the same tag every time the same value appears, so the document stays readable. There is no encryption key and no restore button: the mapping lives only in your browser tab, so keep the original file safe.
A hash or an encrypted string would break that too: the result no longer looks like a name, only a string of characters. A page full of random-looking text is also harder to review before you send it to anyone.
A quick checklist before you choose
- Reading a document as a person? Use consistent labels from a counter or a random generator, with a table you control.
- Matching records across databases without reading them? A keyed hash (HMAC) with a real secret key is far safer than a plain hash.
- Need the exact original value back later? Encryption is the only one of the four built to reverse on purpose.
- Never hash a name, an email address or any value from a short, guessable list without a secret key.
- Whatever you pick, GDPR still treats the result as pseudonymised data, not anonymised data, so your usual duties keep applying.
Frequently asked questions
- Is pseudonymisation the same as anonymisation?
- No. Pseudonymised data can still be linked back to a person with extra information, so GDPR still treats it as personal data. Anonymisation is meant to remove that link for good, which is a much higher bar to clear.
- Can I just hash a name to pseudonymize it?
- You can, but a plain hash of a name or an email is weak on its own. Names and emails come from a guessable list, so an attacker can hash their own list and match the results, as ENISA's 2019 report explains.
- What is the real difference between tokenization and encryption?
- A token is usually a made-up value linked to the original in a secure table, with no maths connecting the two. Encryption calculates the protected value from the original with a key, so that same key can calculate it back.
- Do I need to understand cryptography to pseudonymize a document?
- Not for a document you plan to share and have someone read. Consistent labels from a counter, kept with a table only you control, cover most everyday cases without any cryptography at all.
Sources & references
- Article 4(5) GDPR — definition of pseudonymisation — Intersoft Consulting (GDPR full text)
- Guidelines 01/2025 on Pseudonymisation — European Data Protection Board
- Pseudonymisation Techniques and Best Practices (2019) — European Union Agency for Cybersecurity (ENISA)
- Information Supplement: PCI DSS Tokenization Guidelines — PCI Security Standards Council
Mask a document without uploading it
ONYRI Sanitize finds names, identifiers, bank details and secrets in a PDF, a Word file or a scan, and masks them in your browser. You check the preview, then download a flattened copy.