CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking
Source
arxiv.org
Author
Anjali Kantharuban, Jonas Mueller
Date
Why it matters
Agent benchmarks that use LLM user simulators can misstate success rates and failure modes. Calibrating simulators to real sessions makes tau2-Bench style results better reflect real user behavior.
Key takeaways · AI-distilled
The authors find that existing user simulators match human style but lack outcome calibrationHow well a model's confidence matches reality — a calibrated model saying "90% sure" is right about 90% of the time.Full definition →: they do not reproduce the success rates and failure patterns observed when real users talk to the same AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition →.
CUE encodes observed sessions into continuous embeddingA list of numbers representing a piece of text's meaning, so that similar meanings end up numerically close and can be searched.Full definition →, samples from that space, and decodes samples into persona commands that steer an LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition → to act as the user, with no simulator training required.
On tau-squared-Bench, the paper reports CUE-steered simulators made fewer simulator-caused errors and better matched real-user failure modes, aggregate success rates and per task-user outcomes than other persona-based methods.
Though fit mostly on customer support interactions, the simulators reportedly generalized to document creation, math tutoring and casual conversation, and kept working across different simulator LLMs without refitting CUE.
Terms in this piece · Glossary
embedding — A list of numbers representing a piece of text's meaning, so that similar meanings end up numerically close and can be searched.
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
calibration — How well a model's confidence matches reality — a calibrated model saying "90% sure" is right about 90% of the time.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.