Interpreting and Steering LLM Agents for Social Simulations
Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit. However, LLMs are ultimately black boxes based on deep neural networks which limits their value for social science. This is because of a lack of (i) interpretability: i.e. the ability to assign clear mechanisms driving observed behavior; and a lack of (ii) steerability: i.e. the ability to mute or amplify specific theoretically meaningful mechanisms of action to drive specific model behavior. Here, we demonstrate how the black box could be opened up to further enrich LLM-based simulations. Specifically, we compare three types of methods: (1) prompt-based manipulation, (2) SAE-derived feature steering, and (3) probe-based direction steering and examine their utility for LLM-based social scientific simulations. We do so by interpreting and steering two foundational components of human behaviors, namely preferences (risk attitudes, altruism) and capabilities (divergent creativity, product innovation), operationalized using four classic economic and creative tasks implemented as natural-language interactions.
Introduction. Large Language Models (LLMs) mimic humans on multiple dimensions: they understand natural Having said that, LLM-based simulations might not always faithfully match human behavior Here, we develop a framework to compare alternate methods for looking into the LLM black-box
Discussion / Conclusion. Future work might also explore hybrid approaches, such as multi-probe steering or low-rank
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Do language models reason like humans or mimic surface patterns?- What makes LLM behavior socially interpretable to human observers?
- Can distributional views explain when an LLM appears to change its mind?
- How do different social roles affect LLM theory of mind errors?
- How do LLMs default to surface-level strategies instead of genuine mental simulation?
- Why do users attribute beliefs to LLMs despite uncertainty about their minds?
- Can models track dynamic mental state changes better than static beliefs?
- Can LLMs simulate belief revision in social systems without modeling thought?
- Do realistic LLM behaviors require simulating human thought or just behavior?
- Why does LLM simulation elicit information that direct elicitation cannot?
- Do emotion-driven actions in agent simulators capture genuine belief revision or just reactive behavior?
- How do LLM user simulators fail to represent authentic user behavior distributions?
- What cognitive structures do realistic belief models need to include?
- Do LLMs need world models to make accurate predictions?