
If you're building or evaluating character/persona AI, your fidelity scores may be measuring cultural-composite fluency instead of accuracy — this lays out why and personality benchmarks structurally miss 'Miranda distortion' and offers a concrete three-axis rubric (anachronism detection, documentary consistency, contextual plausibility) to catch it.
“If a dominant failure mode is anachronistic compositing, and your evals measure fluency and personality consistency, then your evals cannot detect the dominant failure.”
Jacob E. Thomas
“the composite Hamilton knows he will be the subject of a Broadway musical. The composite Lincoln has already read the Gettysburg Address, even if he was summoned before he wrote it.”
Jacob E. Thomas
“Compositing is not a bug that you patch in post training. Post training reinforces it.”
Jacob E. Thomas
“You don't train a persona. You compose an encounter. And you keep the receipts.”
Jacob E. Thomas
“You do not want a model that is your mother. You want a model that can speak with your mother's documents in the room.”
Jacob E. Thomas
articleMatrAIx: Simulating the World with 8.3 Billion Persona AgentsXiaomin Li, Yuexing Hao, Jianheng Hou, Jintao Huang, Qianfeng Wen, Shirley Huang, Yifan Liu, Xiaoyi Liu, Yilan Fan, Yijun Wang, Koutian Wu, Ruoqi Gao, Muhammad Ahmed Mohsin, Jing Tang, Brihi Joshi, Heming Liu, Zheyuan Deng, Zonglin Di, Sankalp Jajee, Jiuyao Lu, Zhiwei Zhang, Saksham Kapoor, Ishan Gupta, Yunhan Zhao, Chanwoo Park, Yucheng Lu, Bing Hu, Weihang Xiao, Aravind Mohan, Hanwen Xing, Runyu Zhang, Mihir Kulshreshtha, Yuanda Xu, Qianyu Zhu, Dianzhuo Wang, Yuxin Xiao, Bowen Jiang, Yongye Su, Wenhao Chai, Zuxin Liu, Lawrence Yunliang Chen, Xuandong Zhao, Ethan Ye, Shivam Patel, Jason Xie, Alex Martin Richmond, Weixiang Ding, Emre Okcular, Diya Mathew, Ziheng Wang, Rana M. Shahroz Khan, Zhejian Peng, Fang Wu, Fan Nie, Xinyang Han, Yubin Kim, Jiawei Zhang, Zhenting Qi, Huangyuan Su, Xu Pan, Abinitha Gourabathina, Hyewon Jeong, Hemanth Neelgund Ramesh, Kumail Alhamoud, Kimia Hamidieh, Zidi Xiong, Samuel Schmidgall, Pengrui Han, Yepeng Huang, Yongheng Wang, Bowen Yang, Alex Gu, Yuchu Wang, Akshay Paruchuri, Brenna Li, Hejie Cui, Jiayuan Ding, Chaosheng Dong, Jiahao Wang, Yixuan He, Chi Wang, Pamela Bhattacharya, Tianyi Peng, Paul Pu Liang, Mitchell Gordon, Yilun Du, Marinka Zitnik, James Zou, Prasanna Tambe, Philip Torr, Emily Fox, Asu Ozdaglar, Dawn Song
videoEvaling Video Slop — Maor Bril, Character.aiAI Engineer
articleWhen Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey ResponsesZihan Chen, Di Zhu, Lei Nico ZhengChecking sign-in…
Loading comments…