Google evaluates alignment of behavioral tendencies in large language models
Google Research introduces a framework that converts psychological questionnaires into situational judgment tests (SJTs) to assess behavioral alignment of LLMs. Testing 25 models reveals that smaller models often deviate from human consensus, while larger models show improved but imperfect alignment. The study also finds models are overconfident when human opinions diverge.
Google/DeepMind
