Security and privacy
What is PII masking for AI?
In short
PII masking for AI means finding personally identifiable information, such as names, e-mail addresses, phone numbers, ID numbers and card numbers, and hiding or replacing it before an AI model reads it or before a reply is shown. Detection usually combines patterns and statistical detectors, so it catches most but not all cases. Masking reduces how much personal data reaches the model and its logs, but it works best alongside access permissions, not instead of them.
When an AI assistant answers a question from your documents or tools, everything it reads goes to a language model. That often includes personal data: customer e-mails, phone numbers, employee ID numbers, bank details. PII masking is the practice of finding that data and hiding it before the model sees it, or before a person sees the answer.
PII stands for personally identifiable information: any data that can identify a person on its own or combined with other data. The US NIST SP 800-122 guide to protecting the confidentiality of PII is a good plain reference for what counts and why it matters.
Why does PII matter in AI systems?
AI systems move data in more places than a normal app. A single question can pull text from several documents, call a few tools and send all of it to a model provider. Each step is a chance for personal data to go somewhere it should not.
The OWASP Top 10 for LLM Applications lists sensitive information disclosure (LLM02) as a top risk. Masking reduces the damage in several ways:
- The model provider receives less personal data, so less can end up in its logs.
- The answer is less likely to repeat a phone number or ID number to the wrong person.
- Your own conversation history and logs hold fewer real values.
- It is easier to show auditors that you limit personal data to what a task needs.
How does PII detection work?
Detection comes first. You cannot hide what you have not found. There are three common methods, and most tools combine them.
- Patterns (regular expressions). Rules that match a fixed shape, such as an e-mail address, a US Social Security number or an IBAN. Fast and predictable, but they only find formats you wrote rules for.
- Checksums and validation. Extra checks that cut false alarms, such as the Luhn check that valid card numbers pass.
- Named entity recognition (NER). A model that reads context to spot names, places and organisations. It finds things patterns cannot, but it can miss some and flag others by mistake.
Custom patterns matter as much as built-in ones. Your own customer numbers, policy numbers and internal project codes are sensitive to you, and no generic detector knows their format.
Masking vs redaction vs hashing vs tokenization
These words are often used interchangeably, but they do different things.
| Method | What happens to “[email protected]” | Reversible? | Keeps values distinguishable? | Good for |
|---|---|---|---|---|
| Masking | j***@acme.com | No | Partly | Showing enough for a person to recognise a record |
| Redaction (removal) | [EMAIL] or nothing | No | No | Highest privacy when the value is not needed |
| Hashing | a fixed string such as 3f9a… | No, but guessable for common values | Yes, same input gives the same output | Matching and counting without seeing values |
| Tokenization | a token such as EMAIL_0042, with the real value kept in a separate vault | Yes, by whoever holds the vault | Yes | Restoring real values later for authorised people |
Two notes on the table:
- Hashing is one-way, but short or common values like phone numbers can be recovered by hashing every possible value and comparing. A secret key or salt makes this much harder.
- Tokenization and keyed hashing are forms of pseudonymisation. The GDPR definition in Article 4 describes it as processing so data can no longer be linked to a person without additional information kept separately. Pseudonymised data is still personal data under the GDPR.
Where should masking happen: before the model or after?
There are four points in an AI request where masking can apply.
| Point | What it protects | What it misses |
|---|---|---|
| On the person’s prompt | Data the person types or pastes | Data the AI fetches itself |
| On retrieved documents and tool results, before the model | What the model provider receives and logs | Values the person typed themselves |
| On the model’s output | What the person sees on screen | What the provider already processed |
| On logs and stored history | Data kept after the conversation | Everything that happened in the live request |
Masking before the model is the stronger privacy control, because the provider never receives the real value. Masking only the output is cosmetic from a privacy point of view: the data has already left.
A good setup also decides who may see the real value. A support lead might need a full phone number while everyone else sees a masked one. That rule should follow people and groups, not be hard-coded in each agent.
What are the limits of PII masking?
Be honest about what masking can and cannot do.
- Detection is never perfect. Patterns miss unusual formats, and entity detectors miss some names. Expect some leakage and plan for it.
- Context still identifies people. “The only engineer in our Lisbon office” identifies someone with no PII in it.
- Masking can break a task. If the agent needs the real e-mail address to send a message, you must allow it for that step.
- It is not access control. Masking hides values inside documents a person can read. It does not stop the AI from reading documents the person should never see. For that you need permission-aware retrieval (see permission-aware answers).
How do you set up PII masking for AI?
- List the data classes you hold. E-mail, phone, national ID, card and bank numbers, health notes, plus your own identifiers.
- Decide the action per class. Mask, hash or remove.
- Add custom patterns for your own identifiers.
- Apply masking before the model on documents and tool results, not only on output.
- Name who may see real values, by person or group.
- Test with real samples and review misses regularly.
- Pair masking with permissions and an audit trail.
For the wider picture of how data escapes AI systems, read AI data leakage.
How promptev handles PII masking
- Masking is set per project, with built-in detectors for e-mail, phone, US Social Security numbers, card numbers, IBANs and API keys, plus your own patterns.
- Each match can be masked, hashed or removed.
- Masking applies to found documents and tool results before the model reads them.
- Named people or groups can be allowed to see the real value.
- Masking works alongside permission-aware answers and an audit trail of every run and settings change. See the governance features.
Frequently asked questions
What is the difference between PII masking and PII redaction?
Masking hides part or all of a value but keeps its shape, such as showing only the last four digits of a card. Redaction removes the value completely, often leaving a label like EMAIL in its place. People often use the two words loosely for the same idea.
Can PII detection catch everything?
No. Pattern rules miss unusual formats and detectors that rely on context miss names and addresses written in unexpected ways. Treat masking as a strong layer that lowers exposure, and keep permissions and access controls as the main protection.
Should I mask data before or after the AI model?
Before the model is the stronger choice for privacy, because the model provider never receives the real value. Masking after the model only protects what the person sees, not what the provider processed and may have logged.
Does masking hurt the quality of AI answers?
Sometimes. If the model needs the real value, for example to match a customer by e-mail, masking it removes that ability. Consistent placeholders or hashing keep the ability to tell values apart without revealing them.
Is hashed data still personal data?
Often yes. Under the GDPR, data that can be linked back to a person with additional information is pseudonymised, and pseudonymised data is still personal data. Hashing common values like phone numbers can also be reversed by guessing.
What is AI DLP?
AI data loss prevention (DLP) applies DLP ideas to AI traffic: it inspects prompts, retrieved documents, tool results and outputs for sensitive data, then blocks, masks or logs it. PII masking is one of its main techniques.

Faisal Saeed is Founder & CEO of Promptev, building next-gen context engineering infrastructure that enables teams to orchestrate, scale, and deploy production-ready generative AI systems with confidence.