
It reframes moderation as a placement-and-recovery design problem rather than a classifier-accuracy problem.
“Content-moderation classifiers are usually evaluated in isolation, but deployment requires choosing where to intervene and what follows a flag.”
“At the evaluated operating points, Response only achieves the highest filter-only Usefulness in both settings, while Input + response achieves lower Harmful Exposure.”
“Replacing Response only blocking with Response + rewrite recovers most blocked traffic and yields the same observed Harmful Exposure count as Response only blocking for the selected configuration; this equality is not an equivalence result.”
“These results support comparing moderation configurations under deployment-specific safety and latency constraints rather than applying a universal placement rule.”
Checking sign-in…
Loading comments…