Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media

Paper · arXiv 2412.18148 · Published December 24, 2024
Expertise in the Age of AI Content

Social media platforms are experiencing a growing presence of AI-Generated Texts (AIGTs). However, the misuse of AIGTs could have profound implications for public opinion, such as spreading misinformation and manipulating narratives. Despite its importance, it remains unclear how prevalent AIGTs are on social media. To address this gap, this paper aims to quantify and monitor the AIGTs on online social media platforms. We first collect a dataset (SM-D) with around 2.4M posts from 3 major social media platforms: Medium, Quora, and Reddit. Then, we construct a diverse dataset (AIGTBench) to train and evaluate AIGT detectors. AIGT- Bench combines popular open-source datasets and our AIGT datasets generated from social media texts by 12 LLMs, serving as a benchmark for evaluating mainstream detectors. With this setup, we identify the best-performing detector (OSM- Det). We then apply OSM-Det to SM-D to track AIGTs across social media platforms from January 2022 to October 2024, using the AI Attribution Rate (AAR) as the metric. Specifically, Medium and Quora exhibit marked increases in AAR, rising from 1.77% to 37.03% and 2.06% to 38.95%, respectively. In contrast, Reddit shows slower growth, with AAR increasing from 1.31% to 2.45% over the same period. Our further analysis indicates that AIGTs on social media differ from human-written texts across several dimensions, including linguistic patterns, topic distributions, engagement levels, and the follower distribution of authors. We envision our analysis and findings on AIGTs in social media can shed light on future research in this domain. Our code and dataset are publicly available.1

Introduction. The rapid development of Large Language Models (LLMs) has markedly enhanced the quality of AIGTs, enabling the use of models like GPT-3.5 [42] in daily life to produce high-quality texts, such as in academic writing [14], questionanswering [23], and translation [65]. These AIGTs are often indistinguishable from Human-Written Texts (HWTs), presenting AIGT detection as a crucial yet challenging task for effective classification. On social media platforms, the use of LLMs to answer questions can contribute to the spread of misinformation [72]. Furthermore, AIGTs may be deliberately used for information manipulation or the dissemination of fake news, potentially resulting in serious societal impacts [16]. To better understand the prevalence of AIGTs on social media platforms, we aim to quantify and monitor their presence, addressing the question: On social media, are we already interacting with AI-generated texts? Currently, numerous detectors have been developed to detect AIGTs. According to the MGTBench [17], these detectors are broadly divided into two categories: metricbased [12, 39] and model-based detectors [4, 21, 54], some of which have shown high accuracy and robustness. While these detectors have been applied in controlled settings, recent studies have explored their effectiveness in real-world scenarios. Hanley and Durumeric [16] conduct AIGT detection on news website articles, with a primary focus on content generated by GPT-3.5 and others from Turing benchmark, which includes various pre-2022 models [62]. Furthermore, Liu et al. [31] detects ChatGPT-generated content on arXiv papers. However, academic and news writing are formal and tailored to specific audiences, whereas social media content is more interactive, making it a better domain for observing AIGTs’ impact on daily life. Moreover, previous studies do not account for recent popular LLMs, while we consider a broader range of models in our efforts to detect AIGTs on social media.

To quantify and monitor AIGTs on social media, we collect textual data from 3 popular platforms spanning January 1, 2022, to October 31, 2024, as most LLMs are released after 2022. After data preprocessing, we obtain 1,170,821 posts from Medium, 245,131 answers from Quora, and 982,440 comments from Reddit. We name it as SM-D, short for Social Media Dataset. To identify the most effective detector, we construct a dataset named AIGTBench, which consists of public AIGT/Supervised-Finetuning (SFT) datasets and our own AIGT datasets generated from social media data. AIGTBench includes AIGTs generated by 12 different LLMs, such as GPT Series [44]) and Llama Series [10, 60, 61]), totaling around 28.77M AIGT and 13.55M HWT samples. We then benchmark AIGT detectors on AIGTBench and leverage the bestperforming detector as our primary detector, which achieves an accuracy of 0.979 and an F1-score of 0.980. To better reflect its application in detecting AIGTs on online social media, we rename it as OSM-Det (Online Social Media Detector). Based on OSM-Det, we quantify and monitor the texts across the 3 platforms and use the AI Attribution Rate (AAR) to represent the rate of posts classified as AI-generated (The pipeline is shown in Figure 1). We observe several noteworthy phenomena: (1) A sharp rise in AI-generated content begins in December 2022, with distinct AAR trends emerging across platforms. Before December 2022, the AAR across platforms remains stable. However, starting in December, Medium and Quora show significant surges, while Reddit shows only a slight increase. This suggests the widespread and diverse LLM adoption on social media; (2) Linguistic analysis shows similar AAR trends and exhibits stylistic features in AIGTs/HWTs. Based on the word-level analysis, we find that the usage trend of top-frequency AIpreferred words aligns closely with LLM adoption trends. With sentence-level analysis, we also reveal that AIGTs tend to be more objective and standardized, whereas HWTs are more flexible and informal; (3) Technology-related topics drive higher AARs on Medium. Topics like “Technology” and “Software Development” show the highest AARs, indicating that users with a strong technical background are more likely to adopt LLMs; (4) Predicted HWTs receive more engagement than AIGTs. On Medium, the content predicted as HWTs receives more average “Likes” and “Comments” than AIGTs. This suggests that users are more inclined to engage with HWTs; and (5) Authors with fewer followers are more likely to produce AIGTs. On Medium, users with no more than one thousand followers tend to produce content that has the highest mean AAR at 54.02%. In contrast, as the follower count increases, the AAR gradually shifts toward the lower range (≤25.00%). Our contributions are summarized as follows: • We are the first to conduct a systematic study to quantify, monitor, and analyze AIGTs on social media.

Related work. The growth in model parameters and training data has recently empowered LLMs to demonstrate exceptional language processing capabilities [70]. Since then, LLMs have gradually gained popularity, like GPT-4 [43] and Llama [60], enabling users to generate high-quality texts effortlessly. Yet, LLMs exhibit multiple inherent vulnerabilities [18, 29, 57, 71] and have raised concerns about potential misuse, such as fake news generation [69], academic misconduct [63], hate speech generation [52], and performance degradation of training LLMs using AI content [6], making the detection of AIGTs (also known as machine-generated texts) increasingly important [11]. He et al. [17] introduce MGTBench for standardizing the evaluation of different LLMs and experimental setups within the AIGT detectors. They broadly categorize the detectors into two main types: metric-based and modelbased detectors. Metric-based detectors use pre-defined metrics, such as log-likelihood, to capture the characteristics of texts [12, 39, 56]. In contrast, model-based detectors rely on trained models to distinguish between AIGTs and HWTs [4, 15, 21, 25, 31, 54]. More introduction refer to Appendix B. Besides, some researchers have applied detectors to text detection in real-world scenarios. Hanley and Durumeric [16] train a detector using data generated by the ChatGPT and Turing benchmark model and conduct tests on multiple news websites. Their study reveals that, from January 1, 2022, to May 1, 2023, the proportion of synthetic articles increased on news sites. Liu et al. [31] also conduct detection on arXiv and find a significant rise in the proportion of papers using ChatGPT-generated content, reaching 26.1% by December 2023. In contrast to their detection targets, we focus on detecting AIGTs on social media platforms and covering a broader range of LLMs. Macko et al. [33] construct a multilingual dataset based on instant messaging and social interaction platforms such as Telegram and Discord, using it to compare the performance of existing detectors. In contrast, our research focuses on providing an in-depth temporal analysis of AIGTs on content-driven social platforms like Medium, Quora, and Reddit.

Method. 3 Data Collection In this section, we elaborate on the data collection process, which primarily includes two datasets: the social media dataset (SM-D) and the detector training dataset (AIGTBench).

Unlike previous research, we focus on social media platforms, including Medium, Quora, and Reddit, emphasizing content creation, sharing, and discussion. The introduction of platforms is in Appendix C. These platforms stand out for hosting longer, more detailed posts where users emphasize the depth and quality of the information they share. As shown in Table 1, we collect data from these social media platforms from January 1, 2022 to October 31, 2024. We consider this part as our social media dataset for analysis. For each platform, the detection targets are determined based on their distinct characteristics. On Medium, a blog hosting platform, we extract both the titles and contents of articles, treating the entire article as the detection target. On Quora, a question-and-answer platform, we select the corresponding answers to questions as the detection target. Similarly, on Reddit, which is known for its user-driven discussions, we also choose the response content as the detection target. Furthermore, we apply data filtering with the rules described in Appendix E.

To train the AIGT detectors, we consider two parts of the data. First, we consider 6 publicly available AIGT datasets and 5 common SFT datasets to form the training dataset (see Tables A3 and A4 for dataset statistics and Appendix D for more details). Second, to increase the detector’s generalization capabilities on social media, we additionally collect data from the 3 social media platforms ranging from January 1, 2018, to December 31, 2021, as verified by an ablation study demonstrating that these new subsets fill a gap that older benchmarks missed (see Appendix J). We classify this data as HWTs, given that most LLMs had not been published during this period. We also design different LLMs writing tasks to generate AIGTs that align with the characteristics of platforms (Table A1 describes the statistics details). For Medium, which is primarily used for sharing articles and blogs, the core tasks are centered on writing. We design two LLM writing tasks: (1) polish articles to create polished versions; (2) based on the article’s title and summary, directing the LLM to generate complete article content, thereby simulating a writing scenario. For Quora and Reddit, which mainly focus on question answering and user interaction, we design two tasks: (1) polish texts like Medium and (2) query LLM directly answer questions, simulating a user interaction scenario. Detailed prompts are provided in Appendix F. Overall, the datasets used for training our detector and the distribution of LLM series are shown in Figure 2. This dataset includes 12 different LLMs, with a detailed introduction provided in Appendix A. Within these datasets, the two most prevalent model series are the GPT Series, which accounts for 42.99%, and the Llama series, which represents 39.05%. GPT Series is the most widely used proprietary model and has played a pivotal role in the evolution of generative AI. As of January 2023, approximately 13M users interact daily with GPT-3.5 [66]. The Llama series models also have significant influences, as the report indicates that downloads of Llama models on the Hugging Face platform have nearly reached around 350M [38]. Therefore, these two model series are the primary focus of our dataset. During the data generation process, we notice that certain samples contain textual noise, like irrelevant or redundant information. To maintain data quality, we implement some data processing strategies (see Appendix E for details).

Discussion. demonstrates that OSM-Det exhibits strong generalization capability when detecting AIGTs generated by previously unseen LLMs.

From September 2023 to the first half of 2024, although the AAR remains high, it declines from the peak in early 2023 and gradually stabilizes between 22.03% −30.79% throughout 2024. This indicates that the behavior of Quora users in generating AI content is becoming more stable. From June 2024, the AAR gradually decreases and reaches a low near 19.79% between September and October 2024. The increase in AAR may be attributed to Quora’s launch of its LLM platform, Poe, in 2023 [2], which initially led to a rise in AI-generated content. However, as many Quora users found Poe’s capabilities insufficient to meet their daily needs, the AAR likely declined following this initial surge, eventually stabilizing. while, model-specific analysis shows the words “think” , “Why can’t we restore...” . In summary, the results suggest that sentence-level patterns provide more distinctive characteristics for distinguishing AIGTs and HWTs, as LLMs may usually follow a standardized pattern to generate texts.

We observe a rapid increase in AAR for all topics following the release of GPT-3.5 in December 2022, indicating that the popularity of LLMs has impacted all topics on Medium. Besides, the AAR for “Technology” and “Software Development” remains consistently higher than other topics from December 2022 to October 2024, ranking respectively first and second. One possible reason is that people in the technology field are more likely to know about LLMs and frequently interact with them, leading to a higher AAR.

As shown in Figure A3a, the predicted-AIGTs receive fewer “Likes” on average than predicted-HWTs, with mean values of 69.15 and 127.59, respectively. And predicted- AIGTs exhibit a higher frequency of low “Likes” counts. Figure A3b shows that predicted-AIGTs receive fewer “Comments” on average compared to predicted-HWTs, with mean values of 4.16 and 7.38, respectively. We further investigate the mean values of Likes and Comments for authors with different numbers of followers and Table 5 indicates that, across all follower count groups, AIGTs receive significantly fewer Likes and Comments compared to HWTs. To summarize, predicted-HWTs obtain more “Likes” and “Comments”, which indicates that users in Medium are generally more willing to engage with human-written content. However, the relatively small gap between the two suggests that AI-generated content appeals to users.

Conclusion. In this paper, we collect a large-scale dataset, SM-D, encompassing multiple platforms and diverse time periods, providing the first comprehensive quantification and analysis of AIGTs on online social media. We construct AIGTBench, an AIGT detection benchmark integrating diverse LLMs, to identify the most effective detector, OSM-Det. We then perform temporal tracking analyses, highlighting distinct trends in AAR that are shaped by platform-specific characteristics and the increasing adoption of LLMs. Finally, our analysis uncovers critical differences between AIGTs and HWTs across linguistic patterns, topical features, engagement levels, and the follower distribution of authors. Our findings offer valuable perspectives into the evolving dynamics of AIGTs on social media.

Limitations. In this paper, we conduct long-term quantification of AIGTs on 3 commonly used social media platforms, but there are still some limitations: 1. Limited coverage of LLMs: AIGTBench includes only 12 LLMs and does not cover all LLMs released across different time periods. We only included models released after November 2022. This decision was made because our study specifically focuses on more powerful models, such as ChatGPT, which may lead to misclassifications for earlier models. In addition, the data set shows distributional bias favoring the GPT series 42.9% and the Llama series 39.05% models. While current AIGT detectors can generalize to unseen LLMs to some extent [25], these coverage limitations may introduce slight errors and pose potential impacts on the accuracy of some results. However, these biases are unlikely to significantly impact the analysis results, as these models are also the most widely used in real-world applications. 2. Lack of analysis on multilingual platforms: Our research focuses on English-dominated social media platforms. Therefore, the applicability of our findings is restricted to these specific platforms and language contexts. Since data collection is a long-term process, we plan to gradually expand to multilingual environments and more platforms in future research to improve the universality of the conclusions. 3.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Are AI-generated articles systematically disadvantaged in search ranking and user engagement? How does AI-generated content create social proof without authentic interaction? Does disclosing AI authorship change how audiences evaluate the writing? How reliably can humans and AI detectors identify machine-generated text? How do users confuse explanation quality with actual system accuracy? Can artificial systems establish authority in domains requiring expert judgment?