# Personal data

> Turn on personal data protection, choose what Mindset looks for, add and test your own patterns, and see what was caught.

After this page, you can decide what happens to personal data before it reaches a model provider, add patterns of your own, and check what Mindset has caught. When the invoice exceptions agent reads an invoice with a supplier contact's email and IBAN on it, this is the setting that decides whether the provider sees them.

## What this is

**Personal data protection** finds personal data in what an agent sends to a model and swaps each value for a placeholder before the provider sees it. It's off until an admin turns it on. You set it once for the whole org in **Settings → Governance**, in the **Personal data** card. Only org admins can change it.

Nothing is ever blocked or refused for containing personal data. Detected values are replaced, and the message carries on.

## What is protected, and when

| Question | Answer |
|---|---|
| What's the default? | Off. Messages reach the model provider exactly as written |
| What turns it on? | The org policy, **What happens to personal data in transit**, set to any option other than **Off** |
| Is there a per-connection switch? | Only on AI model connections, labeled **PII protection**. It's on by default |
| What does the per-connection switch do? | It can only remove protection. Turned off, calls through that model connection skip the check. It can't add protection when the org policy is off |
| With the org policy off, does the switch do anything? | No. Nothing is scrubbed anywhere, whatever the switch says |
| Do agent conversations follow the switch? | No. Agent conversations follow the org policy and ignore the switch. The switch applies only to `$llm` calls made through that connection, for example from a function |
| What happens to a detected value? | It's replaced by a placeholder token. Nothing is blocked |
| Are images scrubbed? | No. Images skip scrubbing at every setting, and so does anything a model reads off an image |

Changing the **PII protection** switch asks you to confirm. Turning it off is recorded in the audit log.

## The four policy options

Pick one under **What happens to personal data in transit**, then click **Save**.

| Option | What happens |
|---|---|
| **Off: personal data is not substituted** | Messages reach the provider as written. The default |
| **Redact permanently: the real value never comes back** | The assistant and every tool it calls see placeholders, so anything that needs the real value fails |
| **Redact, with audited access to the real value** | Replaced before the provider sees it, restored on the way back. Each restore is recorded in the audit trail |
| **Redact in transit only** | Replaced before the provider sees it and restored straight away on the way back |

What a placeholder looks like depends on the class. Each one keeps the shape of the original and comes from a range that's never issued to anyone: an email becomes an address ending `@redacted.invalid`, a card number becomes a 16-digit number that still passes the Luhn check, and an IP address becomes one in `192.0.2.x`. Values your own patterns catch become a short marker instead.

## Choose what Mindset looks for

Under **What to look for** there are six built-in classes, each with its own switch. All are on by default.

| Class | How it's recognized |
|---|---|
| **Email** | Shape match, no checksum available |
| **Phone** | Parsed as a number, not shape-matched |
| **Card** | 13 to 19 digits passing the Luhn checksum |
| **IBAN** | Country code and check digits, mod-97 validated |
| **National ID** | Structurally validated (US SSN today) |
| **IP address** | Octet-range checked v4, conservative v6 |

A class you switch off is still detected, then discarded before substitution. It isn't replaced, and it isn't counted while it's off.

## Add and test your own pattern

Under **Your own patterns**, you add a regular expression for something the built-in classes don't know, such as a supplier account number on your invoices.

1. Enter a **Name**, for example "Supplier account".
2. Enter the **Pattern**: a regular expression, without delimiters or flags.
3. Paste some sample text into **Test against**. It isn't stored.
4. Click **Test pattern**. The result shows how many matches it found and how long the pattern took.
5. Click **Save pattern**. It stays disabled until a test passes and the pattern has a name.

Each saved pattern has its own on and off switch and a **Remove** button.

Limits, because every active pattern runs on every model call:

- At most 25 patterns can be on at once. You can keep more saved and switched off.
- A pattern can be at most 200 characters.
- A pattern that can match an empty string is refused.
- A pattern that takes longer than 25 ms on the test text can't be saved. Nested repetition is the usual cause.

## See what was caught

The detections feed is in **Govern → Audit**, on the **Personal data** tab. The Settings card links there with **See what these rules have caught, in Audit**.

Under **What was caught**, each row shows **When**, **Class**, **Placeholder**, **Found by** and **Restored**. The real value is never shown there, at any policy setting, to anyone. See [Audit](https://docs4.mindset.ai/docs/ams/audit) for the rest of that tab.

## Send telemetry to your own collector

**Telemetry export (OpenTelemetry)** sits on the same **Settings → Governance** tab. It pushes your agents' traces and metrics to Datadog, Grafana, Honeycomb or any OTLP collector.

1. Enter the **OTLP endpoint**: the base URL of your OTLP/HTTP collector or vendor intake.
2. Enter the **Auth header name** your backend expects (it starts as `Authorization`) and the **Auth token**. The token is stored as a reference and never shown again.
3. Choose what to send: **Send traces**, **Send metrics** and **Send agent behavior (script phases and gate outcomes)** start on. **Send costs** starts off. **Include end-user identifiers** starts on, and puts the person's email or user ID on every span.
4. Click **Test**. **Save** is enabled only after the endpoint accepts the test telemetry. Mindset tests again when you save and refuses the save if that test fails.

Message content isn't exported. The only personal identifier in the export is the end-user identifier, controlled by its own switch. Under **Is it flowing?** you see recent batches and whether each one was delivered.

## What can go wrong

- **The model connection says PII protection is on, but nothing is replaced.** The org policy is **Off**. The connection switch can't turn protection on by itself.
- **A tool fails after you choose Redact permanently.** The tool received a placeholder where it needed the real value. Use **Redact, with audited access to the real value** or **Redact in transit only** if tools need real values back.
- **Personal data in a screenshot or scanned invoice reached the provider.** Images skip scrubbing. Keep images with personal data out of agents that send them to a model. See [What Mindset does not do](https://docs4.mindset.ai/docs/ams/what-mindset-does-not-do).
- **Your pattern won't save.** It hasn't passed a test, it's too slow, it matches an empty string, or 25 patterns are already on.
