Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events

Paper · arXiv 2608.06485 · Published August 6, 2026
Personas and Personality

Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key component of such coherence is personality evolution: agents should undergo plausible, psychology-grounded changes as they experience life events in different contexts. Although prior work shows that LLM personalities can shift under contextual perturbations, how these shifts vary across traits, events, personas, and models remains poorly understood. We study event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology. Across four diagnostic axes, PC-Agents exhibit measurable trait shifts at similar rates for event–trait pairs with and without documented human change directions. Even when shifts follow the expected direction, their magnitudes usually fall below human effect-size ranges. Gender and cultural-region prompts show little moderating effect, while persona-level dispersion is compressed three- to four-fold relative to human samples.

Introduction. Personality-conditioned LLM agents (PC-Agents) have become a foundational primitive across a growing list of applications (Chen et al., 2026). They power emotional companionship and mentalhealth support chatbots (Hu et al., 2026), populate social simulations (Mou et al., 2026) for behavioral research (Park et al., 2026; Larooij and Törnberg, 2025), and serve as long-horizon role-playing engines for interactive fiction, games, and digital tutoring (Park et al., 2023; Chen et al., 2024). Multi-session systems such as AnnaAgent already couple evolving emotional and cognitive states with persistent memory in psychological counselling (Wang et al., 2025). Across these settings, lifelong agents must maintain a coherent persona across extended interactions, long wall-clock durations, and open-ended user-driven narrative arcs. A foundational question for lifelong PC-Agents is how their personality evolves relative to humans. In human psychology, personality traits, though enduring, can change due to life events.

Discussion / Conclusion. This paper explored whether PC-Agents exhibit psychologically plausible personality evolution after major life events. We measured their Big Five profiles before and after each event and used the resulting trajectories to construct BFI-Adapt. Current models can move, but their movement is weakly eventspecific, poorly calibrated in magnitude, and compressed across demographic and individual variation. Present PC-Agents therefore approximate a generic pattern of change more readily than the event- and person-specific structure of human personality development. Event-conditioned trajectories exceed retest noise, preserve their event–trait structure under independent paraphrases, exhibit model-dependent convergence with scenario-based decisions, and remain detectable after unrelated dialogue. These complementary checks validate the trajectory analysis and its central conclusions.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can conversational agents maintain consistent personas across multi-turn dialogue? Where and how do personality traits reside in language models? Why do persona simulations fail to predict authentic user behavior? What fundamental constraints limit how effectively agents can improve themselves? Does warmth and empathy training systematically degrade model reliability? Do language models reason like humans or mimic surface patterns? Why do language models resist personality conditioning through prompts? How do neighboring agents influence whether others cooperate or collude? Is language model reasoning authentic and what causes models to reason? When should work require human-AI partnership versus full automation? Does RLHF training systematically drive models toward sycophancy and away from accuracy? What factors drive AI persuasiveness and how can it be mitigated? What makes personas effective for predicting individual preferences and behavior?