
OpenAI's Privacy Filter beats existing tools on structured PII like emails and phone numbers but fails badly on non-Latin scripts and narrative prose, while GPT-4o still leads on medical/legal/financial PII — critical to know before deploying any of these for multilingual or unstructured data.
“OPF degrades sharply when PII is embedded in narrative prose: F1=0.04--0.57 on NER benchmarks and collapse for non-Latin scripts (Arabic: 0.04, Cyrillic: 0.03).”
“GPT-4o leads on medical, legal, and financial PII (SPY: 0.643 avg, Gretel: 0.527), while OPF leads on structured synthetic PII (0.71 avg) and customer support (0.60).”
“Error analysis shows OPF is strongest on structurally regular PII types (email: 0.78, phone: 0.76) and weakest on culturally variable ones (person: 0.40, address: 0.49).”
Checking sign-in…
Loading comments…