Machines in the Crowd? Measuring the Footprint of Machine-Generated Text on Reddit

Paper · arXiv 2510.07226 · Published October 8, 2025
Expertise in the Age of AI Content

Generative Artificial Intelligence is reshaping online communication by enabling large-scale production of Machine-Generated Text (MGT) at low cost. While its presence is rapidly growing across the Web, little is known about how MGT integrates into social media environments. In this paper, we present the first large-scale characterization of MGT on Reddit. Using a state-of-the-art statistical method for detection of MGT, we analyze over two years of activity (2022–2024) across 51 subreddits representative of Reddit’s main community types such as information seeking, social support, and discussion. We study the concentration of MGT across communities and over time, and compared MGT to human-authored text in terms of social signals it expresses and engagement it receives. Our very conservative estimate of MGT prevalence indicates that synthetic text is marginally present on Reddit, but it can reach peaks of up to 9% in some communities in some months. MGT is unevenly distributed across communities, more prevalent in subreddits focused on technical knowledge and social support, and often concentrated in the activity of a small fraction of users. MGT also conveys distinct social signals of warmth and status giving typical of language of AI assistants. Despite these stylistic differences, MGT achieves engagement levels comparable than human-authored content and in a few cases even higher, suggesting that AI-generated text is becoming an organic component of online social discourse. This work offers the first perspective on the MGT footprint on Reddit, paving the way for new investigations involving platform governance, detection strategies, and community dynamics.

Introduction. Generative Artificial Intelligence (GenAI) is transforming the Web and reshaping everyday practices related to search [40], news consumption [32], entertainment [2], and online social interactions [4]. In particular, the dynamics of the participatory Web are being revolutionized by the commercial release of Large Language Models (LLMs) that dramatically lowered the barriers to producing Machine-Generated Text (MGT). The broad diffusion of these new tools is paving the way for large volumes of crowd-generated synthetic content to be uploaded to social media and other participatory online platforms. This socio-technical shift carries profound societal implications. The Web—and especially social media—has become a central arena for public discourse and information sharing, shaping the spread of ideas and influencing democratic participation. The diffusion of MGT may significantly affect these collective processes [7] in ways that scholars are still actively debating. Optimistic perspectives emphasize the potential of GenAI to enhance contributors’ creativity [15], increase engagement, and promote inclusivity in online discussions [37]. In contrast, critics raise concerns about misinformation, declining authenticity in user interactions [17, 37],deteriorating content quality [17, 30], and emergent biases [3]. When produced with strategic or malicious intent, MGT can be deliberately crafted to be persuasive [8] and highly engaging [33], while embedding inaccurate, biased, or harmful information [17]. This could exacerbate systemic risks identified by the European Digital Services Act, including discrimination, declining mental well-being, and the erosion of civic and electoral processes [38]. Developing a quantitative understanding of AI’s impact on people’s online experiences is therefore both essential and urgent to inform effective policies and design safeguards against potential harms. Although distinguishing Machine-Generated Text (MGT) from Human-Generated Text (HGT) text is inherently challenging, promising tools have been developed with high accuracy and inference speed [5]. These tools have enabled initial estimates of the proliferation of synthetic content online [36] and of the growing presence of AI agents posing as human users [33]. However, beyond these initial estimates, the use of MGT in real-world online social environments remains mostly unexplored. To shed light on how online social dynamics might change as result of MGT, it is not only important to estimate its incidence, but also to analyze the nature of the content produced, the context in which it is published, and the response that the public has to it. To help address this gap, we present the first large-scale characterization of MGT on Reddit, one of the world’s leading social media platforms. With its millions of active users and countless topic-based discussions, Reddit stands as a compelling case in point for exploring whether and to what extent MGT emerges and spreads with human discourse online. Focusing on 51 popular subreddits representative of Reddit’s main functional community types, we analyze comments and submissions from 2022 to 2024, a period that includes milestones in the release of GenAI tools. We used a state-of-the-art method to detect messages with a high likelihood of being MGT and address four main research questions:

RQ1 — How is MGT adopted across subreddit communities and over time, and how does this reflect the distinct conversational norms of different community categories? RQ2 — What temporal patterns characterize MGT prevalence, and how might these relate to exogenous events? RQ3 — How does the distribution of MGT users evolve over time, and are there any significant trends in its adoption? RQ4 — Do MGT and HGT differ in the type or intensity of social signals and engagement patterns they convey across subreddit categories?

Related work. The growing presence of MGT across online platforms has motivated a growing body of research aimed at estimating its prevalence and understanding its potential impact on digital communication. Studies using BERT-based detectors on 15M news articles revealed that by mid-2023, since the release of ChatGPT the proportion of AIgenerated news increased by over 57% on mainstream outlets and by an astonishing 474% on misinformation and propaganda sites [21]. Similarly, supervised detectors estimated that AI-generated articles account for roughly 40% of news contributions on Medium and Quora [36]. Analyses of Wikipedia entries based on the GPTZero detector suggest that over 5% of newly created English articles may be AI-generated [9]. Even the scientific review process appears affected: recent estimates indicate that between 6% and 17% of peer reviews for leading Machine Learning conferences may be produced by LLMs [25]. In contrast, fewer studies have quantified the prevalence of AIgenerated text on social media. Earlier research primarily detected machine-generated posts by identifying bot-like authors exhibiting coordinated or automated behaviors [13]. Text-based approaches relying on stylometry [23] offered limited accuracy and generalizability. More recent studies have leveraged crowdsourced signals (such as X’s Community Notes) to examine the diffusion and nature of AI-generated discourse [16]. On Reddit, community responses to AI-generated content have been characterized by skepticism and concern [27], leading moderators and administrators to introduce new governance rules soon after the public release of ChatGPT [26]. Crowdsourced data indicate that AI-generated visual content remained relatively rare through late 2023, representing fewer than 0.5% of image-based posts [28]. In this context, the most closely related study to ours is the recent work by Sun et al. [36], who applied a supervised detection method to 982K Reddit posts published between January 2022 and July 2024, estimating that around 2.5% were AI-generated. However, their analysis focuses solely on prevalence and does not investigate the nature, context, or engagement dynamics of machine-generated content across communities—gaps that our work seeks to address.

Method. To characterize the emergence and diffusion of AI-generated texts on Reddit, we start by considering the top 1,000 subreddits by number of subscribers.1 The choice of analyzing large and active communities is motivated by (i) the broad visibility and potential impact that the content posted in those communities have on the public, and (ii) the substantial volume of data providing sufficient support for a reliable estimation of the prevalence of MGT. We manually parsed the list of subreddits in decreasing order of popularity and selected a representative subset of 51 subreddits that we mapped into five main functional subreddits categories that have been informed by a taxonomy form prior work [39]: • Information Seeking, including, e.g., r/worldnews, r/askscience, r/explainlikeimfive, r/health; • Social Support, including, e.g., r/GetMotivated, r/mentalhealth, r/AITAH, r/LifeProTips; • Discussion, including, e.g., r/changemyview, r/unpopularopinion, r/SeriousConversation, r/politics; • Identity, including, e.g., r/teenagers, r/asktransgender, r/AskWomen, r/BlackPeopleTwitter; • Chit Chat, including, e.g., r/funny, r/entertainment, r/books, r/Showerthoughts. We collected the complete data on submissions and comments posted in these subreddits from January 2022 to December 2024 (see Appendix A.1 for the complete list of subreddits) using data dumps obtained from the PushShift API [6]. This covers the period during which LLMs and GenAI have become widely accessible to the public, allowing us to study their adoption and diffusion in online discourse through Reddit. After filtering out empty posts (e.g., containing only images or URLs), we collected 38,074,021 comments and 4,073,586 submissions. A detailed overview of the number of analyzed comments for each month is reported in Figure A1, Appendix A.3.

Problem setting. We are given a collection of text messages (e.g., Reddit comments or submissions) X = {x1, ...,xn} where each message xican be authored by either a human (i.e., HGT) or any machine-generator tool like large language models (i.e., MGT), but the type of author is unknown. We frame the detection of MGT content in online discussions as a binary text classification task: given a classifier model f, the goal is to assign a label yi∈{0, 1} to each xi∈Xsuch that:

Approach. To detect signals of generative AI usage within Reddit communities, we considered two major classes of MGT detection techniques: • Metric-based approaches [5, 18, 29, 34, 35] rely on statistical properties of texts such as token distributions, entropy measures, perplexity; • Model-based approaches [10, 19, 20, 34] leverage deep-learning classifiers trained for distinguishing HGT from MGT. Although both approaches have been widely adopted in the literature and demonstrated strong performance [22], we opted for a metric-based solution because of two main practical constraints. First, model-based approaches are typically trained on datasets from domains like news articles or essays, leading to degraded performance on conversational data such as Reddit comments. Moreover, fine-tuning such approaches on Reddit data would be costly in practice, as it would require a large training set of labeled Reddit-specific data—currently unavailable, to the best of our knowledge. Second, model-based approaches require relatively slow (neural) inference for each input text, which becomes computationally prohibitive at a scale of millions of comments and submissions. Based on these considerations, we selected the zero-shot metricbased Fast-DetectGPT [5], as our primary detection tool. Fast-Detect- GPT strikes a good balance between detection capabilities and computational efficiency, ranking among the most effective detectors in recent benchmarks [24] while remaining significantly faster than model-based detectors and commercial alternatives (e.g., GPTZero).

Discussion. With this first large-scale study of Machine-Generated Text (MGT) on Reddit, we contribute to advancing the understanding of how Generative AI is reshaping discussions and information sharing in online social media. We moved beyond estimating the overall prevalence of MGT and instead examined its concentration and nature within an online ecosystem. Our results show that different communities are affected to varying degrees, with MGT more concentrated in subreddits oriented toward technical knowledge exchange and social support (RQ1).

After the public release of ChatGPT, the use of MGT has been steady over time, with peaks of prevalence corresponding to major releases of new LLMs (RQ2). Most of the MGT is concentrated in the activity of a few users, with 2% of the active users being responsible for the entire MGT production (RQ3). We also found MGT often conveys social signals that differ from those of HGT. Most notably, across many subreddits, MGT exhibits strong patterns of social support and status-giving, reflecting the characteristic style of modern AI assistants. Finally, MGT comments are often indistinguishable from HGT in terms of positive engagement, and in some cases even outperform them (RQ4).

Impact. These findings have several implications. Practically, they highlight the need for platforms to adapt moderation and transparency policies, particularly in communities where trust and authenticity are central, such as knowledge-sharing and support fora. Theoretically, they suggest that GenAI is not merely amplifying existing discourse but actively shaping new communicative norms, challenging established notions of authenticity and prompting to consider how human and AI voices co-evolve in online ecosystems. Ethically, the stronger engagement elicited by MGT opens questions about the risks of manipulation and representational imbalance.

Limitations. Our contribution has limitations that future work could address. First, we study a single platform, making the findings not generalizable to other social media. Reddit is well-suited for this type of analysis due to its relatively unconstrained comment length, which facilitates MGT detection. Extending the analysis to platforms where textual content is typically much shorter (e.g., X, Bluesky) would require the development of specialized MGT detectors for short-form text—a research challenge in itself. Second, while our selection of subreddits includes popular, active, and diverse communities, it still represents a small portion of the activity on Reddit. These communities are therefore not necessarily representative of the full spectrum of GenAI use or MGT diffusion across the platform. Nevertheless, by demonstrating that the spread of MGT on Reddit is non-negligible, our study highlights the need for broader, more systematic investigations. Third, although informative temporal patterns emerge from our analysis, the public adoption of GenAI tools is still too recent to draw any conclusions about long-term trends of MGT use. Within our observation window, MGT prevalence appears relatively stable after the initial surge of adoption between November 2022 and September 2023.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How reliably can humans and AI detectors identify machine-generated text? How does AI-generated content create social proof without authentic interaction? How can AI systems reliably guide voters without introducing political bias? Are AI-generated articles systematically disadvantaged in search ranking and user engagement? How do users confuse explanation quality with actual system accuracy? Can artificial systems establish authority in domains requiring expert judgment? Can AI systems participate in genuine communication or only simulate it? How does tokenization reshape what we value in intelligence?