phi-redact
Finds patient names, MRNs, dates, phone numbers and 26 other kinds of PHI and PII in English text - on the device, so it never reaches your servers, your logs or an LLM provider.
11 million parameters, 11 MB as int8. Small enough to run in the browser tab you're reading this in, which is exactly what the demo below does.
Redact as you type
Pick a sample, paste your own text, or dictate it. Every sample is invented - but anything you paste or say stays in this tab either way.
Restore map - stays on the device
Loading model…
How well it works
Strict span-level scoring: a prediction counts only if both the character boundaries and the label match.
0.860
Micro F1 on Nemotron-PII
13,134 entities, strict span match
86.1%
Recall, 5-benchmark average
Desert Ant Redact (EN): 86.5%
11.1M
Parameters
BERT-mini, 4 layers
11.3 MB
int8 ONNX in your browser
~5.6 MB est. as 4-bit Core ML
Against Desert Ant Redact v0.4.0 (English)
Desert Ant figures are their published English results on the same datasets.
| phi-redact v7 | Redact (EN) | |
|---|---|---|
| WikiANN name recall | 84.3% | 69.5% |
| MultiNERD name recall | 94.3% | 94.9% |
| Composite recall (5 benchmarks) | 86.1% | 86.5% |
| Parameters | 11.13M | 23M |
| Est. 4-bit on-device size | ~5.6 MB | 11.6 MB |
| Languages | English | 27 |
| License | Apache-2.0 | Source-available |
F1 by entity, Nemotron-PII
2,000 test documents. PROFESSION over-fires on job titles in ordinary prose.
- MRN 0.951
- ADDRESS 0.949
- NAME_PATIENT 0.944
- NAME_FAMILY 0.937
- LICENSE 0.920
- COUNTRY 0.919
- SSN 0.908
- IDNUM 0.908
- EMAIL 0.905
- CREDITCARD 0.905
- CITY 0.904
- ROUTING_NUMBER 0.876
- DATE 0.859
- URL 0.828
- DEVICE 0.791
- PROFESSION 0.494
Where it fits
Server-side redaction means the raw text has already crossed the network to get redacted. Doing it on the device removes that hop.
Before the LLM call
Swap names, MRNs and phone numbers for placeholders like [NAME_PATIENT_1], send the prompt, then put the originals back in the reply. The model provider never sees the patient.
Intake and case forms
Scrub free-text fields in the browser before submit, so the notes box on a benefits or intake form doesn't quietly become the place PHI ends up in your database.
Logs and analytics
Run it at ingest on support transcripts, error reports and session notes. Nothing leaves the machine to be redacted, so there's no second system holding the raw text.
Test data from real data
Mask a production sample into something a developer can safely use to reproduce a bug - consistent placeholders keep the same person the same person across a record.
30 entity types
Grouped the way the demo colours them.
Names
NAME_PATIENTNAME_FAMILYNAME_PROVIDER Contact
PHONEFAXEMAILURLIP Location
ADDRESSCITYSTATEZIPCOUNTRY Dates & age
DATEAGE Medical
MRNHEALTHPLANHOSPITALDEVICE Government IDs
SSNLICENSEPASSPORTIDNUMVIN Financial
ACCOUNTCREDITCARDIBANROUTING_NUMBER Work
EMPLOYERPROFESSION Get started
The same weights ship as safetensors, ONNX, Core ML and LiteRT in the model repo.
redact-core.js is the tokenizer, windowing and span decoder this page runs - one dependency-free file.
Get it on GitHub ↗
npm install onnxruntime-web import * as ort from "onnxruntime-web";
import { loadRedactor, redact, restore } from "./redact-core.js";
const redactor = await loadRedactor({
ort,
baseUrl: "https://huggingface.co/CDT5058/phi-redact/resolve/main/",
});
const input = "Patient Maria Garcia, DOB 1985-03-22, MRN 4471829.";
const spans = await redactor.detect(input);
// [{ start: 8, end: 20, label: "NAME_PATIENT", text: "Maria Garcia", score: 0.9999 }, ...]
const { text, map } = redact(input, spans);
// "Patient [NAME_PATIENT_1], DOB [DATE_1], [MRN_1]."
const reply = await callYourLLM(text); // the provider only sees placeholders
console.log(restore(reply, map)); // originals put back locally pip install onnxruntime tokenizers huggingface_hub numpy import json
import numpy as np
import onnxruntime as ort
from huggingface_hub import hf_hub_download
from tokenizers import Tokenizer
repo = "CDT5058/phi-redact"
session = ort.InferenceSession(hf_hub_download(repo, "phi-redact-v7-int8.onnx"))
tokenizer = Tokenizer.from_file(hf_hub_download(repo, "tokenizer.json"))
tags = json.load(open(hf_hub_download(repo, "tags.json")))
text = "Patient Maria Garcia, DOB 1985-03-22, MRN 4471829, called 415-555-0189."
enc = tokenizer.encode(text) # adds [CLS]/[SEP], truncates at 256 tokens
feed = {
"input_ids": np.array([enc.ids], dtype=np.int64),
"attention_mask": np.array([enc.attention_mask], dtype=np.int64),
"token_type_ids": np.array([enc.type_ids], dtype=np.int64),
}
pred = session.run(["logits"], feed)[0][0].argmax(-1)
spans, start = [], None
for (s, e), tag in zip(enc.offsets, (tags[i] for i in pred)):
if tag == "O" or s == e:
start = None
continue
pos, label = tag[0], tag[2:]
if pos in "BS":
start = s
if pos in "ES" and start is not None:
spans.append((label, text[start:e]))
start = None
print(spans)
# [('NAME_PATIENT', 'Maria Garcia'), ('DATE', '1985-03-22'),
# ('MRN', 'MRN 4471829'), ('PHONE', '415-555-0189')] from transformers import AutoTokenizer, AutoModelForTokenClassification
import torch
tokenizer = AutoTokenizer.from_pretrained("CDT5058/phi-redact")
model = AutoModelForTokenClassification.from_pretrained("CDT5058/phi-redact").eval()
text = "Patient Maria Garcia, DOB 1985-03-22, MRN 4471829, called 415-555-0189."
enc = tokenizer(text, return_tensors="pt", truncation=True, max_length=256,
return_offsets_mapping=True)
offsets = enc.pop("offset_mapping")[0].tolist()
with torch.no_grad():
pred = model(**enc).logits.argmax(-1)[0].tolist()
tags = [model.config.id2label[i] for i in pred]
# Decode BIOES tags to character spans exactly as in the ONNX example. Specs
- Task
- Token classification, BIOES tags (121 labels: 30 entity types × 4 + O)
- Architecture
- BERT-mini (google/bert_uncased_L-4_H-256_A-4), 4 layers, 256 hidden, 4 heads
- Training
- Distilled (α = 0.7) from a 66M-parameter DistilBERT teacher on ~157k commercial-safe documents - no real patient data
- Context
- 256 tokens per pass; longer text runs as overlapping windows
- Language
- English only - other languages miss silently rather than degrade safely
- Formats
- safetensors (44.5 MB fp32) · ONNX fp32 (44.6 MB) and int8 (11.3 MB) · Core ML · LiteRT
- License
- Apache-2.0 weights; training data CC-BY 4.0 / CC-BY-SA
Know the limits
Not a certified de-identification system.
It doesn't satisfy HIPAA Expert Determination or Safe Harbor. Treat it as a detection aid alongside human review.
- English only. Other languages produce silent misses.
- Trained on synthetic and public data - typos, OCR noise and unusual formatting will cost accuracy.
- URL recall drops to 0.318 in informal first-person prose.
- Roughly 1 in 7 entities is missed on Nemotron-PII. Don't treat a clean output as proof there's nothing left.
Questions
Does the text I type here get sent anywhere?
No. The page downloads the model once (11.3 MB, served from this site) and ONNX Runtime - fetched from jsDelivr - runs it in WebAssembly in your tab. There's no inference endpoint - you can open the network tab and watch nothing go out while you type.
Where does dictation happen?
Also on the device. The first time you press Dictate, the page downloads Whisper tiny (about 41 MB) and transformers.js; after that your recording is transcribed in WebAssembly and handed straight to the redactor. The audio is never uploaded - it's the same arrangement as the text.
Is this HIPAA de-identification?
No. It isn't certified under Expert Determination or Safe Harbor. It's a detection aid: use it to reduce exposure and to pre-mark text for review, and evaluate it on your own documents before relying on it anywhere compliance-sensitive.
How does it compare to Desert Ant's Redact?
On English it's roughly tied - 86.1% vs 86.5% average recall across five benchmarks - at half the parameters and about half the on-device size. Redact covers 27 languages; phi-redact is English only and adds clinical types like MRN, health plan and provider names.
Why are some labels wrong in the demo?
The model sometimes gets the right span with the wrong type - an insurer tagged as EMPLOYER, or the last four digits of a card as a ZIP code. For redaction the span matters more than the label, but it's worth knowing. PROFESSION is the weakest type (F1 0.494) because it fires on job titles in ordinary prose; switch it off if you don't need it.
What happens with long documents?
The model reads 256 tokens at a time. The demo slides overlapping windows across the text and keeps each token's prediction from the window where it sits closest to the middle, so entities at a boundary still get full context.
Putting redaction in front of an LLM, or working with PHI in a regulated system? Reach out.