Mapping the Emerging Social Science of Large Language Models

Paper · arXiv 2609.07598 · Published September 7, 2026
Therapy Practice and AI

Large language models (LLMs) have moved from specialized text-generation systems into everyday and institutional settings, where they increasingly shape communication, learning, work, creativity, and decision-making. Research on these developments has grown rapidly across disciplines and publication venues, but it remains fragmented and lacks an integrated framework for organizing the field. This study maps the emerging social science of LLMs through a curated corpus of 198 papers reviewed in full and a field-scale corpus of 47,719 formally published papers retrieved from five bibliographic databases. We combine sentence embeddings, K-means clustering, within-cluster Latent Dirichlet Allocation (LDA), author and LLM classifications, and structural topic modeling to identify and validate the field’s organization. The analyses recover three domains: LLM as Social Minds, concerning socially interpretable model behavior; LLM Societies, concerning collective dynamics among interacting model-based agents; and LLM–Human Interactions, concerning how people perceive, use, and are affected by LLMs.

Introduction. Large language models (LLMs) have moved rapidly from text-generation systems to interfaces through which people seek advice, learn, create, collaborate, and make decisions. The same We use the term social science of LLMs to describe systematic research that treats LLMs or LLM-based agents themselves as social objects of explanation [199]. The field is defined by what a study seeks to explain: socially meaningful model behavior, interactions involving This study develops and evaluates such a taxonomy. It addresses three questions. First, what We answer these questions through two complementary studies. Study 1 analyzes a curated corpus of 198 papers reviewed in full by the authors. Titles and abstracts are represented with MPNet sentence embeddings and partitioned using K-means solutions across K= 2–9, with internal validation and stability analyses used to assess alternative resolutions.

Discussion / Conclusion. This study examined how research on the social science of LLMs can be organized across two complementary corpus scales. In Study 1, the selected three-cluster solution was stable under resampling and supported a substantive interpretation in terms of LLM as Social Minds, LLM Societies, and LLM–Human Interactions. The correspondence of this solution with both the authors’ full-text classifications and the LLM-based title-and-abstract classifications showed that these distinctions could be applied through different evaluative procedures, while the The conceptual contribution of this framework begins with a shared object of explanation: LLMs or LLM-based agents themselves, examined through their behaviour, interactions, and social consequences. Within this common scope, the three domains distinguish the principal relation requiring explanation. LLM as Social Minds focuses on socially interpretable capacities and These findings also suggest a research agenda centred on explaining the mechanisms that connect Several limitations delimit the present map. The Study 1 corpus was purposively curated and does not provide exhaustive coverage, although Study 2 enabled its coverage and broader

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What safeguards enable trustworthy AI-assisted scientific peer review at scale? Is language model reasoning authentic and what causes models to reason? Do language models respond to social pressure and face-saving like humans? Does encoded knowledge in language models actually influence their outputs? Do language models reason like humans or mimic surface patterns? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? Why don't LLMs reliably translate capability into accurate outputs? Why does memory consolidation cause performance regression in continual learning? How do prompt design choices influence model reasoning and performance? How do LLM judges' systematic biases affect alignment and evaluation outcomes? Why do LLM recommenders underperform collaborative filtering despite their capabilities? Can multi-agent systems avoid converging on false agreement without deliberation? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? How do false presuppositions and sycophancy drive persistent false beliefs in models?