What if the way we collect human feedback in robotics is quietly losing information?
If one trajectory drops the cutlery and another drops the plate, asking for a single preference hides information behind the choice.
Freeform Preference Learning lets annotators describe preference axes in natural language, learns rewards conditioned on those axes, and extracts better policies.
Congrats to @marceltornev, @anubhamahajan01,@AbhijnyaBhat, and @chelseabfinn!
Standard binary preference labels hide information (e.g. why one robot trajectory was preferred); this method recovers that information to train more informative reward models.