
Substituting predictions for user responses forfeits the randomization guarantee; identification then rests on unverifiable surrogacy assumptions that weaken as a treatment departs from past experiments. Useful before trusting any synthetic-user .
articleWhen Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey ResponsesZihan Chen, Di Zhu, Lei Nico Zheng
articleA/B test models in productionwww.together.ai
articleHow well LLM-based test generation techniques perform with newer LLM versions?Michael Konstantinou, Renzo Degiovanni, Mike PapadakisChecking sign-in…
Loading comments…