Vibeleaderboard
← All Intel
Intel / article

CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking

Source
arxiv.org
Author
Anjali Kantharuban, Jonas Mueller
Date
Why it matters

Agent benchmarks that use LLM user simulators can misstate success rates and failure modes. Calibrating simulators to real sessions makes tau2-Bench style results better reflect real user behavior.

Key takeaways · AI-distilled
  • The authors find that existing user simulators match human style but lack outcome : they do not reproduce the success rates and failure patterns observed when real users talk to the same .
  • CUE encodes observed sessions into continuous , samples from that space, and decodes samples into persona commands that steer an to act as the user, with no simulator training required.
  • On tau-squared-Bench, the paper reports CUE-steered simulators made fewer simulator-caused errors and better matched real-user failure modes, aggregate success rates and per task-user outcomes than other persona-based methods.
  • Though fit mostly on customer support interactions, the simulators reportedly generalized to document creation, math tutoring and casual conversation, and kept working across different simulator LLMs without refitting CUE.
Terms in this piece · Glossary
  • embedding — A list of numbers representing a piece of text's meaning, so that similar meanings end up numerically close and can be searched.
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
  • calibration — How well a model's confidence matches reality — a calibrated model saying "90% sure" is right about 90% of the time.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Recommended reads
Comments

Checking sign-in…

Loading comments…