ResearchAgents 🇺🇸 27.07.2026 16:04

ConvApparel: Measuring and Closing the Realism Gap in LLM-Based User Simulators

Google/DeepMindGoogle/DeepMind
Google Research introduces ConvApparel, a new human-AI conversation dataset and three-pillar evaluation framework to quantify and reduce the 'realism gap' in LLM-based user simulators. The framework includes population-level statistics, human-likeness scoring, and counterfactual validation, applied to simulators built with Gemini models. Results show data-driven simulators (ICL, SFT) outperform prompt-based ones, but even the best models still exhibit subtle synthetic artifacts.
Researchers at Google Research, Ofer Meshi and Sally Goldman, introduced ConvApparel, a new dataset and evaluation framework designed to measure and improve the realism of LLM-based user simulators. The dataset comprises over 4,000 human-AI multi-turn conversations (nearly 15,000 turns) in the apparel shopping domain, collected using a dual-agent protocol where users were randomly routed to either a helpful 'Good' agent or an intentionally unhelpful 'Bad' agent. The framework includes three pillars: population-level statistics, human-likeness scoring via an automated discriminator, and counterfactual validation. Using the Gemini model family, three simulators were tested: a prompted baseline, an in-context learning (ICL) simulator with retrieval-augmented generation, and a supervised fine-tuning (SFT) simulator. Results showed that data-driven simulators (ICL and SFT) outperformed the prompted baseline, closely mirroring human behavior in population-level tests, but all simulated conversations were confidently identified as synthetic by the human-likeness discriminator. Counterfactual validation demonstrated that ICL and SFT simulators could adapt to the unseen 'bad' agent by showing increased frustration, while the prompted baseline remained unnaturally polite. The research highlights the persistent realism gap and provides tools to bridge it.
Сокращения
ICL = in-context learning
SFT = supervised fine-tuning
Source: Google Research — original
Our earlier posts on this topic ↓
Fresh news