← All IntelClip / EducationMeasuring alignment via distribution comparison, not right/wrong scoring
From Synthetic Personas for Market Research: Where LLM Agents Break · ≈16:43
“There are many ways for distributions to get wrong.”
“So that sets a noise floor as how accurate our models could ever get because the humans themselves are fundamentally noisy.”
“the key smart thing they did is they took those humans and they brought them back 2 weeks later and they redid the battery of surveys and personality tests and they found that the humans on average were only 80% consistent to themselves.”
What’s in it
- Explains how to score AI persona simulations against noisy human survey data
- Shows why distribution shape matters more than just matching the average
- Gives a workaround to estimate noise floor without retesting real humans
Clip transcript
going to improve your statistical significance for the most part. So what you need to do is you need to do what you do with weather forecast. You'd basically check against what actually happened or in our case what humans actually said. And that's where we're going to basically be measuring distributions of data. Unlike classic e-vals where there's clearly a right and wrong and you can score how many were right and how many wrong, now we need to measure the data as a comparison of distributions. And there are many ways for distributions to get wrong. They could be completely wildly off. They can as we mentioned get the average right but the shape of the distribution wrong. And so you're going to need multiple metrics to capture how well your model is reflecting different personas. Um I recommend using a correlation type metric along with one of these shape type metrics which capture what the underlying shape of the distribution is. The other thing you need to do is estimate the fundamental noise in your ground truth data. That experiment I talked about in the beginning where they got 83% accuracy, the key smart thing they did is they took those humans and they brought them back 2 weeks later and they redid the battery of surveys and personality tests and they found that the humans on average were only 80% consistent to themselves. So that sets a noise floor as how accurate our models could ever get because the humans themselves are fundamentally noisy. And so the 83% is actually normalized against that. If you can do this and bring your humans back, that's great. Very often you can't. So the way you can kind of artificially do this is take your ground truth human data, break it into two chunks, and then pretend one is synthetic and one is human, and then measure the correlation and repeat that hundreds and thousands of times and average it, and that'll set kind of a noise floor that your ground truth data where half of it's synthetic, half of it real, could be the level of accuracy you could hope to get. So hopefully by now you have an
Recommended reads
- clipOperational QC failures and synthetic data issuesAI Engineer
- clipDisagreement as signal, not noise, in human QAAI Engineer
articleMatrAIx: Simulating the World with 8.3 Billion Persona AgentsXiaomin Li, Yuexing Hao, Jianheng Hou, Jintao Huang, Qianfeng Wen, Shirley Huang, Yifan Liu, Xiaoyi Liu, Yilan Fan, Yijun Wang, Koutian Wu, Ruoqi Gao, Muhammad Ahmed Mohsin, Jing Tang, Brihi Joshi, Heming Liu, Zheyuan Deng, Zonglin Di, Sankalp Jajee, Jiuyao Lu, Zhiwei Zhang, Saksham Kapoor, Ishan Gupta, Yunhan Zhao, Chanwoo Park, Yucheng Lu, Bing Hu, Weihang Xiao, Aravind Mohan, Hanwen Xing, Runyu Zhang, Mihir Kulshreshtha, Yuanda Xu, Qianyu Zhu, Dianzhuo Wang, Yuxin Xiao, Bowen Jiang, Yongye Su, Wenhao Chai, Zuxin Liu, Lawrence Yunliang Chen, Xuandong Zhao, Ethan Ye, Shivam Patel, Jason Xie, Alex Martin Richmond, Weixiang Ding, Emre Okcular, Diya Mathew, Ziheng Wang, Rana M. Shahroz Khan, Zhejian Peng, Fang Wu, Fan Nie, Xinyang Han, Yubin Kim, Jiawei Zhang, Zhenting Qi, Huangyuan Su, Xu Pan, Abinitha Gourabathina, Hyewon Jeong, Hemanth Neelgund Ramesh, Kumail Alhamoud, Kimia Hamidieh, Zidi Xiong, Samuel Schmidgall, Pengrui Han, Yepeng Huang, Yongheng Wang, Bowen Yang, Alex Gu, Yuchu Wang, Akshay Paruchuri, Brenna Li, Hejie Cui, Jiayuan Ding, Chaosheng Dong, Jiahao Wang, Yixuan He, Chi Wang, Pamela Bhattacharya, Tianyi Peng, Paul Pu Liang, Mitchell Gordon, Yilun Du, Marinka Zitnik, James Zou, Prasanna Tambe, Philip Torr, Emily Fox, Asu Ozdaglar, Dawn Song
Comments
Sign in to comment.
Loading comments…