We are releasing Swisscoding Name Filter, a family of models focused on detecting personal names in clinical documents. The release includes individual models for English, German, French, and Italian, alongside one multilingual model covering all four languages.
In our local evaluation on the MultiGraSCCo dataset, the multilingual model achieved 97.92% precision, 98.76% recall, and 98.34% F1 across 252 document versions.
The goal is practical: create a name detection model that works on realistic clinical documents and is small enough to be deployed within a hospital's own CPU-based infrastructure.
01 Motivation
The open-source PII landscape can look almost solved from public leaderboards. Many recent models report strong results across dozens of PII categories, often across multiple languages.
In our evaluation on realistic multilingual clinical documentation, many of these models had insufficient recall for personal names to be suitable for our intended use.
We think two things are happening.
- First, much of the public training and evaluation data is synthetic in ways that make the task easier than production de-identification. The model learns formatting artifacts, placeholder conventions, and repeated templates instead of learning to distinguish a person's name from the surrounding clinical text.
- Second, there is no free lunch in a single model that detects 50 or more PII categories across many languages. Detecting telephone numbers, social security numbers, addresses, and many other categories may compete with recognizing ambiguous personal names for the model's limited capacity.
So we started with one question: what if we made a small model that does one thing well?
02 Model scope
Swisscoding Name Filter is a bidirectional token classifier with two labels:
| Label | Meaning |
|---|---|
| PERSON | A token that is part of a private person's name |
| O | Any token that is not part of a private person's name |
The model is based on ModernBERT-base, which has 22 layers and about 149 million parameters. The model can run locally on ordinary CPU servers. Sensitive text does not need to leave the hospital environment just to be screened for names.
03 Training approach
1. Start with de-identified real clinical documents
We use 30,000 medical documents in German, French, and Italian from which personal information has already been removed.
2. Insert structured placeholders
We insert placeholders for personal information at contextually appropriate positions with Qwen3.5-122B-A10B. These placeholders identify the locations where synthetic values will be inserted, enabling labels to be derived from the replacements.
3. Translate into four languages
We generate English, German, French, and Italian versions using Qwen3.5-122B-A10B, choosing the translation direction based on the source language. This produces approximately 30,000 examples per language.
4. Expand each document with multiple personas
For each document, we replace the placeholders with different synthetic personas. The inserted names vary: surnames alone, full names with or without middle names, names with initials, and names with spelling errors. We also account for related information, such as names in email addresses.
| Stage | Illustrative text |
|---|---|
| Placeholder | {{PATIENT_NAME}} was referred by Dr. {{CLINICIAN_NAME}}. |
| Persona A | Nora Keller was referred by Dr. Elias Meier. |
| Persona B | Moretti was referred by Dr. Romano. |
| Persona C | Camille Laurent was referred by Dr. Martin. |
Why we did not rely on fully synthetic source documents
Fully synthetic documents are often too simple and repetitive to represent the complexity of clinical records. Our inspection also found substantial data and annotation quality problems, illustrated by the four traceable examples below.
We inspected source_text, target_text, and privacy_mask in ai4privacy's PII-Masking-300k dataset at commit c8c77895a005822682b66ab547fc0422579bc1d3, and text and spans in NVIDIA's Nemotron-PII dataset at commit b70ffaf5ff39e079776134c5bf4381f00a9fd1ed. The four records below were selected to illustrate distinct failure modes. Each quote is an exact fragment of the corresponding source text. Expand each example to inspect its quoted source text and relevant labels. For ai4privacy, each displayed ID is the value of that record's id field in PII-Masking-300k. The source file and line used to verify it are in the expanded evidence.
Four traceable dataset examples
A missing-value marker is labeled as a name.
The source says Heir N/A N/A. The two N/A spans, [174,177) and [178,181), are labeled GIVENNAME1 and LASTNAME1.
Inspect cited sample and labels
source_text excerpt: - **Name:** Heir N/A N/A privacy_mask: {"value": "N/A", "start": 174, "end": 177, "label": "GIVENNAME1"} {"value": "N/A", "start": 178, "end": 181, "label": "LASTNAME1"}Download complete source file · data/validation/1english_openpii_8k.jsonl · line 51 (one-based)
A person's name receives a birth-date label.
The source lists Mathangi followed by the birth date 13th June 2004. Both Mathangi at [263,271) and the date at [298,312) are labeled BOD in privacy_mask.
Inspect cited sample and labels
source_text excerpt: 2. **Mathangi** - **Date of Birth:** 13th June 2004 privacy_mask: {"value": "Mathangi", "start": 263, "end": 271, "label": "BOD"} {"value": "13th June 2004", "start": 298, "end": 312, "label": "BOD"}Download complete source file · data/train/1english_openpii_30k.jsonl · line 23,297 (one-based)
Name spans overlap and cross a JSON field boundary.
The source includes "full_name":"Noëlle Shanuga Henseler". Its GIVENNAME1 span [160,187) covers the full name and trailing JSON punctuation, while the stored annotation value contains GIVENNAME2_A(, which is absent from that source span. The GIVENNAME2 and LASTNAME1 spans overlap; the latter reaches into the next sex field.
Inspect cited sample and labels
source_text excerpt: "full_name":"Noëlle Shanuga Henseler", "sex":"F" privacy_mask: {"value": "Noëlle GIVENNAME2_A(Shanuga", "start": 160, "end": 187, "label": "GIVENNAME1"} {"value": "Shanuga LASTNAME1_A(Henseler", "start": 167, "end": 195, "label": "GIVENNAME2"} {"value": "Henseler\",\n \"sex\":\"SEX_A(F", "start": 175, "end": 215, "label": "LASTNAME1"}Download complete source file · data/train/french_openpii_31k.jsonl · line 2,952 (one-based)
An unresolved field token is labeled as a first name.
The source contains the literal token medical_record_number. The entire span [346,367) is labeled first_name, although it is a field marker rather than a person's name.
Inspect cited sample and labels
text excerpt: The medical_record_number is required to verify my identity. spans: {"start": 346, "end": 367, "text": "medical_record_number", "label": "first_name"}
The Nemotron viewer link opens the current dataset revision. The UID and pinned commit above identify the record we inspected.
04 Evaluation
We evaluated name detection primarily on MultiGraSCCo, which is derived from extensively altered clinical texts. Its annotations originate from manually labeled German source documents. We also evaluated the models on Nemotron PII as an additional synthetic benchmark.
All scores below are percentages. The tables report Precision, Recall, and F1 where available.
MultiGraSCCo: multilingual results
Combined German, French, Italian, and English results across 252 document versions.
| Model | Precision | Recall | F1 |
|---|---|---|---|
| Swisscoding Name Filter · multilingual | 97.92 | 98.76 | 98.34 |
| OpenMed SuperClinical Large | 93.85 | 84.36 | 88.85 |
| OpenAI Privacy Filter | 79.47 | 71.19 | 75.10 |
For OpenMed, we combined token predictions for first_name and last_name into one name category before calculating Precision, Recall, and F1. We evaluated the MultiGraSCCo scores for OpenMed and Privacy Filter locally; they are not official results published by the respective organizations.
MultiGraSCCo: results by language
Each table uses all available documents in the named language and includes the corresponding monolingual model, the multilingual model, and the baseline models. The Swisscoding models use a 0.5 decision threshold.
English
| Model | Precision | Recall | F1 |
|---|---|---|---|
| Swisscoding Name Filter · English | 100.00 | 99.72 | 99.86 |
| Swisscoding Name Filter · multilingual | 100.00 | 99.16 | 99.58 |
| OpenMed SuperClinical Large | 98.39 | 94.82 | 96.57 |
| OpenAI Privacy Filter | 97.73 | 90.15 | 93.79 |
German
| Model | Precision | Recall | F1 |
|---|---|---|---|
| Swisscoding Name Filter · German | 97.20 | 99.10 | 98.14 |
| Swisscoding Name Filter · multilingual | 96.92 | 99.40 | 98.14 |
| OpenMed German SuperClinical Large | 98.07 | 56.36 | 71.58 |
| OpenMed SuperClinical Large | 93.13 | 75.35 | 83.31 |
| OpenAI Privacy Filter | 68.35 | 66.03 | 67.17 |
French
| Model | Precision | Recall | F1 |
|---|---|---|---|
| Swisscoding Name Filter · French | 99.35 | 97.68 | 98.51 |
| Swisscoding Name Filter · multilingual | 97.43 | 97.04 | 97.24 |
| OpenMed French SuperClinical Large | 97.21 | 44.59 | 61.13 |
| OpenMed SuperClinical Large | 93.30 | 85.91 | 89.45 |
| OpenAI Privacy Filter | 78.87 | 67.37 | 72.67 |
Italian
| Model | Precision | Recall | F1 |
|---|---|---|---|
| Swisscoding Name Filter · Italian | 96.90 | 99.59 | 98.23 |
| Swisscoding Name Filter · multilingual | 98.31 | 99.59 | 98.95 |
| OpenMed Italian SuperClinical Large | 95.22 | 35.14 | 51.34 |
| OpenMed SuperClinical Large | 91.65 | 86.43 | 88.97 |
| OpenAI Privacy Filter | 81.16 | 68.55 | 74.32 |
For OpenMed, we combined token predictions for first_name and last_name into one name category before calculating Precision, Recall, and F1. We evaluated the MultiGraSCCo scores for OpenMed and Privacy Filter locally; they are not official results published by the respective organizations.
Nemotron PII: evaluation results
We evaluated the English and multilingual Swisscoding Name Filter models on the complete Nemotron PII test split: 100,000 records, comprising 50,000 US and 50,000 international records. NVIDIA describes Nemotron PII as an English-language dataset; US and international refer to locale conventions, not document languages. The multilingual model's Nemotron score therefore measures performance on English text. Only the Swisscoding rows below come from our evaluation of the full test split.
| Model | Precision | Recall | F1 |
|---|---|---|---|
| Swisscoding Name Filter · English | 94.56 | 99.70 | 97.06 |
| Swisscoding Name Filter · multilingual | 93.94 | 99.58 | 96.68 |
| OpenMed · first_name | 99.48 | 99.51 | 99.50 |
| OpenMed · last_name | 99.42 | 99.29 | 99.35 |
The Swisscoding rows are from our local evaluation of token-level name detection on the full Nemotron PII test split. Each token is classified as PERSON (part of a private person's name) or O (other text). Precision, Recall, and F1 use a 0.5 threshold. This table does not include span-level scores.
The OpenMed first_name and last_name values come directly from OpenMed's published evaluation results, not our 100,000-record run. We did not evaluate OpenMed on our full Nemotron test split. Its evaluation set and label definitions differ from ours, so these numbers are presented only as external reference results.
05 References
- Swisscoding Technologies, model release pages (model details and downloads): multilingual, English, German, French, and Italian
- OpenAI, Introducing OpenAI Privacy Filter
- Answer.AI, ModernBERT-base model card
- Warner et al., Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
- Baroud et al., MultiGraSCCo: A Multilingual Anonymization Benchmark with Annotations of Personal Identifiers
- Edin et al., Symphony for Medical Coding: A Next-Generation Agentic System for Scalable and Explainable Medical Coding
- ai4privacy, PII-Masking-300k dataset snapshot and dataset license
- NVIDIA, Nemotron-PII dataset snapshot (CC BY 4.0)
- OpenMed, OpenMed PII per-label evaluation results