New Research Examines Large Language Models’ Ability to Simulate Human Behavior and Personality
Multiple recent studies explore how large language models can simulate human behavioral patterns, with varying approaches and findings.
According to research published on arxiv.org, open-weight LLMs were evaluated on their ability to simulate human survey populations using a cross-instrument calibration task. The study conditioned personas on real respondents’ verbatim answers to one psychometric instrument and measured them on a second instrument, validated against a 2,058-person human panel. Across a 139-pair grid, simulated cross-instrument correlations tracked real human correlations at r = 0.70-0.73 across all tested models. However, the research found that “this capability does not improve monotonically across model releases: on a matched panel, the newest of three tested Llama releases performs worst on two of three headline metrics.”
A separate study on arxiv.org examined whether Prospect Theory adequately describes LLM decision-making behavior. The researchers found that “PT does not consistently provide a reliable account of LLM decision-making across models, and that its application to LLMs is likely sensitive to epistemic uncertainty.”
In related work on personality modeling, researchers on arxiv.org proposed a mechanistic interpretability approach using sparse autoencoders to identify latent directions corresponding to OCEAN personality traits. According to the study, the method demonstrated promise “in controllably steering personality traits at the mechanistic level while maintaining high performance on standard benchmarks.”