
Any giving advice or mediating from a single user's narrative risks quietly endorsing a one-sided account. This names the failure mode and gives a to measure whether a model asks for the other side before judging.
articleAgentic Scaffolding Amplifies Sycophantic Behavior in Large Language ModelsThantham Jittham
articleSelf- and Other-Labels Induce Bidirectional Bias in LLM JudgesSongeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
articleAgreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical JudgmentsOctavian M. Machidon, Alina L. Machidon, Vojko Strahovnik, Mateja Centa Strahovnik, Jonas Miklav\v{c}i\v{c}, Marko Robnik \v{S}ikonjaChecking sign-in…
Loading comments…