
The same machinery that suppresses harmful output is general-purpose output control — worth understanding if you ship or moderation layers that a downstream operator could repurpose.
articleAgreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical JudgmentsOctavian M. Machidon, Alina L. Machidon, Vojko Strahovnik, Mateja Centa Strahovnik, Jonas Miklav\v{c}i\v{c}, Marko Robnik \v{S}ikonja
articleAI Evaluation Should Work With HumansJan Kulveit, Gavin Leech, Tom\'a\v{s} Gaven\v{c}iak, Raymond Douglas
articleAI Agents Push Humans Out of the LoopMargaret Mitchell, Avijit Ghosh, Samir PassiChecking sign-in…
Loading comments…