If you're building hallucination detection or other classifiers with limited labeled data, this shows how to bootstrap from out-of-domain, permissively-licensed datasets to cut annotation costs while still hitting task performance.
How to use open-source, permissive-use data and collect less labeled samples for our tasks.
Transcript
How to use open-source, permissive-use data and collect less labeled samples for our tasks.