# Scrubbing PII before it leaves your process, not after it reaches Sentry

> Server-side scrubbing runs after the data has crossed the network. beforeSend runs in your process, and it is the only layer you fully control.

Someone opens an issue in Sentry to debug a failed checkout and finds the
customer's email, their delivery address, the full `Cookie` header, and (in
the request body) the raw contents of a form that included a phone number.

Nobody added any of that. Sentry collected it, because collecting request
context is what an error tracker is for, and the defaults are tuned for
debugging rather than for data protection.

## Why the dashboard setting is not enough

Sentry has server-side data scrubbing, and it works. But look at where it runs:
your process serialises the event, sends it over the network to Sentry's ingest
endpoint, and *then* the scrubbing rules are applied before storage.

So by the time server-side scrubbing acts:

- the data has left your infrastructure and your jurisdiction
- it exists in transit logs and possibly in ingest buffers
- if your rules do not match the field, it is stored

The only scrubbing you fully control is the kind that happens in your process,
before the request is made. That is `dataCollection`, `beforeSend`,
`beforeBreadcrumb` and `beforeSendSpan`. Use the dashboard settings as a second
line, never as the first.

## The wrong way

```ts
Sentry.init({
  dsn: process.env.NEXT_PUBLIC_SENTRY_DSN,
  // no dataCollection: Sentry 11's defaults collect cookies, bodies and more
  beforeSend(event) {
    if (event.user) delete event.user.email;   // the one field you thought of
    return event;
  },
});
```

Two failures. Left at its defaults, Sentry 11's `dataCollection` attaches
cookies, request and response bodies, query strings, database values and local
variables, and fills in the user from the request. (Sentry 10 and earlier had a
`sendDefaultPii` flag for this; 11 replaced it with `dataCollection`.) And the
scrubber is an allowlist of one: it handles `user.email` and misses
`request.data`, breadcrumb payloads, tags, and any email that happens to be
inside an exception message.

The deeper problem is that the approach does not survive time. Someone adds a
field to a form next quarter, and nothing about this code notices.

## The right way: deny by default

```ts
Sentry.init({
  dsn: process.env.NEXT_PUBLIC_SENTRY_DSN,
  dataCollection: {
    userInfo: false,
    cookies: false,
    httpBodies: [],
    urlQueryParams: false,
    databaseQueryData: false,
    stackFrameVariables: false,
  },
  beforeSend: scrubErrorEvent,
  beforeBreadcrumb: scrubBreadcrumb,
  // Sentry 11 streams spans one by one; beforeSendTransaction no longer runs.
  beforeSendSpan: scrubSpan,
});
```

with a scrubber built on four rules:

**1. Drop request bodies. Do not filter them.**

```ts
if (event.request) {
  const { data: _data, cookies: _cookies, headers, query_string: _qs, url, ...rest } = event.request;
  event.request = { ...rest, url: redactUrl(url), headers: redactHeaders(headers) };
}
```

A body is arbitrary user content, and "we redact the fields we thought of"
stops being true the moment someone adds a field. If a body were genuinely
needed to debug, you would want a specific, named, minimal projection of it,
not the whole thing minus a blocklist.

**2. Reduce the user to an opaque id.**

```ts
event.user = event.user?.id ? { id: String(event.user.id) } : {};
```

The user field answers one question ("how many people are affected") and an
opaque id answers it completely. An email answers it no better and turns your
error tracker into a copy of your user table, in a third-party system, with a
wider access list than your database.

**3. Redact by value shape, everywhere.**

Key-based rules miss anything inside a string. `` throw new Error(`No user for ${email}`) ``
puts an email in the issue *title*, where it is indexed and searchable, and no
field-level rule will catch it. So run patterns over every string in the event:
emails, long digit runs, bearer tokens, provider key prefixes, JWTs.

**4. Drop console breadcrumbs entirely.**

Breadcrumbs are collected automatically, and console breadcrumbs carry whatever
was logged. A single `console.log(user)` left in from a debugging session
becomes a user record attached to every error for the next hour.

```ts
if (breadcrumb.category === "console") return null;
```

You lose some debugging convenience. You gain the guarantee that a stray log
statement cannot leak a record.

## Do not forget the URL

URLs are the quiet leak. Query strings carry tokens (`?reset_token=...`),
emails (`?email=...`) and search terms people typed. Path segments carry ids.

Strip the query string entirely and replace id-shaped path segments:

```
/orders/9f3c1a2b-…/items?email=ada@example.com   →   /orders/:id/items
```

This has a second benefit: it improves grouping, because a hundred distinct
URLs collapse into one route.

## Make the scrubber total

It runs inside error handling. If it throws, you lose the error you were trying
to report *and* the error from the scrubber is thrown inside the SDK, where it
may be swallowed. So:

- depth-limit the walk (a deeply nested object should not recurse forever)
- guard against cycles with a `WeakSet`
- truncate long strings before running regexes over them
- no non-null assertions, no unguarded property access

Bound your regexes too. An unanchored greedy pattern run over every string of
every event on a busy app is a measurable CPU cost in the hot path of your
error handler.

## What to keep

Scrubbing that removes everything makes the tracker useless, and people
respond by turning it off. Keep what is diagnostic and impersonal:

- HTTP method, route pattern, status code
- an opaque user id, and a plan or role tag
- browser, OS, release, environment
- counts, durations, enums, booleans
- your own fingerprint hint

That set answers "what broke, for how many people, since which deploy", which
is every question triage actually asks.

## Verify with a real event

Configuration review is not verification. Trigger an error on a route that
receives a realistic request, open the event, and read every section: Request,
User, Tags, Contexts, Breadcrumbs, and the exception message itself. Do it
again whenever you add a form, a header or a new integration.

Then write the patterns you added into a unit test, so the next refactor cannot
quietly remove them.

---

Agentic Boilerplate: A Next.js repo your agent already knows. Free during launch, then $99 once.

- Site map for agents: https://agenticboilerplate.com/llms.txt
- Public API: https://agenticboilerplate.com/openapi.json
- Contact: agenticstudio@gmail.com
