ConvApparel: Measuring and Closing the Realism Gap in LLM-Based User Simulators
Google Research introduces ConvApparel, a new human-AI conversation dataset and three-pillar evaluation framework to quantify and reduce the 'realism gap' in LLM-based user simulators. The framework includes population-level statistics, human-likeness scoring, and counterfactual validation, applied to simulators built with Gemini models. Results show data-driven simulators (ICL, SFT) outperform prompt-based ones, but even the best models still exhibit subtle synthetic artifacts.
Google/DeepMind
