In a nutshell
Detection of personal data runs inside the same system as everything else — no additional processor that gets to see your matter texts in the clear, and no model left to guess names.
Why this is a question at all
Anyone buying an AI cloud for their firm faces a second problem: before a text goes to a model, the personal data in it should disappear. One common route is to purchase a dedicated detection service for that. It solves the technical question — and creates a new one: that service has to see the text before masking, that is, in the clear. For your duty of care in selecting providers this means one more place that processes matter secrets, one more contract, one more entry in your record of processing activities.
What we do instead: look it up rather than guess
A detection model has to infer from the sentence that "Kraus" is a person here. We do not have to infer it — we already know. Who is client, opposing side, witness or expert in your matter is already typed in your fact graph; it was created while reading the documents. When masking, we look it up there instead of estimating — including the inflected forms that a plain word list fails on (genitive, short form, adjectival derivation).
On top of that comes what can be recognized structurally and needs no model at all: legal forms such as GmbH, AG or GmbH & Co. KG are recognizable by their suffix, IBANs and ID numbers by their checksum, salutation and diagnosis patterns by their fixed shape.
Three things that follow
A rule set is verifiable. You can read up on why a piece of data was replaced. A model cannot answer that question — it decided, but it cannot justify its decision. For a documented selection decision, that is the difference between evidence and an assertion.
Quality is measured, not estimated. We test detection against a fixed body of 246 test documents with 280 marked spans. The measurement is part of our build check: if the detection rate falls below a defined floor, the build fails. We use the same body for the counter-test — 22 documents deliberately containing terms that must not be masked: courts, authorities, statutory provisions. That, too, now has a hard floor: not a single false alarm.
No further contract. There is no provider you would have to entrust your matter texts to for this step.
What we expressly do not claim
No masking method catches everything, and we do not say otherwise. We name two limits openly, because they follow from the design:
We structurally do not catch courts and authorities. Our organization detection works off the legal form — "Amtsgericht München" does not carry one. That is not an oversight but the very mechanism that keeps the zero-false-alarm commitment: a rule wide enough to catch authority names would inevitably pick up statute names and place names as well.
Legal terms stay put. "Erwerbsminderungsrente" (reduced-earning-capacity pension) is a legal term, not a health statement about a particular person — we leave it untouched. These are exactly the cases a detection model regularly confuses.
And the most important caveat remains the professional one: removing names alone does not make a text anonymous if the attribution follows from the context. That is why masking is one protective measure among several here — not the only one, and not a guarantee.
Where you can see this
The switch is in your firm settings; which agents mask and which deliberately work with clear names is listed there individually. The details are in the handbook under data protection and masking, the context on our compliance page and in the provider comparison.






