Tortured phrases: A dubious writing style emerging in science. Evidence of critical issues affecting established journals

Paper · arXiv 2107.06751 · Published July 12, 2021
Domain Specialization in LLMs

Abstract Probabilistic text generators have been used to produce fake scientific papers for more than a decade. Such nonsensical papers are easily detected by both human and machine. Now more complex AI-powered generation techniques produce texts indistinguishable from that of humans and the generation of scientific texts from a few keywords has been documented. Our study introduces the concept of tortured phrases: unexpected weird phrases in lieu of established ones, such as ‘counterfeit consciousness’ instead of ‘artificial intelligence.’ We combed the literature for tortured phrases and study one reputable journal where these concentrated en masse. Hypothesising the use of advanced language models we ran a detector on the abstracts of recent articles of this journal and on several control sets. The pairwise comparisons reveal a concentration of abstracts flagged as ‘synthetic’ in the journal. We also highlight irregularities in its operation, such as abrupt changes in editorial timelines. We substantiate our call for investigation by analysing several individual dubious articles, stressing questionable features: tortured writing style, citation of non-existent literature, and unacknowledged image reuse. Surprisingly, some websites offer to rewrite texts for free, generating gobbledegook full of tortured phrases. We believe some authors used rewritten texts to pad their manuscripts. We wish to raise the awareness on publications containing such questionable AI-generated or rewritten texts that passed (poor) peer review. Deception with synthetic texts threatens the integrity of the scientific literature.

Introduction. In science there is a history of scholarly publishing stings (Faulkes, 2021). Scholars and journalists have submitted nonsensical papers to various venues to expose dysfunctional peer review. These nonsensical papers submitted can be written by humans (e.g., the Sokal Affair and Bohannon, 2013) or computer generated (e.g., SCIgen, Mathgen). Computer programs designed to generate fake papers and sting publishers are also reused by academic tricksters who easily produce the (fake) publications or (fake) citations they desperately need. As a result, meaningless randomly generated scientific papers end up being served and sometimes sold by various publishers with a prevalence estimated to 4.29 papers every one million papers (Cabanac & Labbé, in press; Van Noorden, 2021). Such papers can be easily spotted by both human and machine; natural language generation tools thus appear to be a cheap and dirty alternative to buying publications from paper mills, which also seems on the rise (Else & Van Noorden, 2021; Mallapaty, 2020). The major recent advances in language models based on neural networks may sooner or later lead to a new kind of scientific writing. Incorrigible optimists would consider that automatic translation, writing enhancement, and summarising tools help authors to produce better scientific papers. Whole books are now generated from thousands of articles used as input (Beta Writer, 2019; Day, 2019; Visconti, 2021). But the generative power of modern language models can also be considered a threat to the integrity of the scientific literature. For example, the dangerous nature of the GPT-3 language model (Brown et al., 2020) was discussed extensively (Hutson, 2021). With this in mind, we report observations about a reputable journal along several lines: occurrences of tortured phrases in publications (e.g., ‘flag to clamor’ in lieu of the established ‘signal to noise’), indication – if not evidence – of AI-generated abstracts, as well as questionable texts and images (including reuse from other sources without proper acknowledgement), as well as recent changes in editorial management (including shortened time between reception and acceptance of manuscripts). Without any definitive proof, we thus provide hints of the rise of a new kind of probably synthetic, nonsensical scientific texts. The outline of this open call for investigation is as follows. Section 2 reports a set of ‘tortured phrases’ spotted in the literature. We then focus our study on Microprocessors and Microsystems, an Elsevier journal in which they concentrate (Sect. 3). We report intriguing irregularities in the editorial timelines of this journal (Sect. 4). The presence of synthetic text generated by advanced language model is hypothesised and Sect. 5 reports the screening of recent publications using an off-the-shelf software detecting synthetic text. Section 6 provides factual evidence of inappropriate and/or poor quality publications. We discuss possible sources of synthetic papers in Sect. 7 before concluding with a call to the scientific community for further investigation on this matter (Sect. 8).

Method. While reviewing recent publications, we encountered an unusual and disappointing phenomenon: well-known and well-established scientific terms were replaced by unconventional phrases. In a typical case, a word-by-word synonymical substitution is applied to a multi-word term. We call tortured phrases these phrases that are incorrectly used in lieu of well-established ones. Table 1 shows some tortured phrases that we were able to find in the literature (at first by chance and then by snowballing with already identified terms) and retro-engineer to infer the correct wording that readers would expect.

On May 25, 2021 we queried the Dimensions academic search engine (Herzog, Hook, & Konkiel, 2020) to retrieve the set of papers containing tortured phrases known at that date (see Fig. 1). Note that some tortured phases may be used in a legitimate way (e.g., ‘enormous information’ in certain contexts) and that the full-text indexing performed by Dimensions ignores punctuation. This may lead to retrieve few articles not using a tortured phrase. Dimensions was chosen for its coverage of the literature that is larger than the Web of Science and Scopus (Singh, Singh, Karmakar, Leta, & Mayr, 2021) and because it is free for scientometric research.1 Founded in 1976, the Microprocessors journal2 was quickly renamed Microprocessors and Microsystems starting from Volume 3 in 1978. It is now published by Elsevier3 and classified by Scopus4 in four subject areas of Computer Science:

In what follows, we conduct a more in-depth analysis of this venue over the period February 2018 to June 2021 for which we collected data. Figure 2 shows a radical change in the number of articles published per volumes starting in 2020.

Microprocessors and Microsystems publishes articles with DOIs minted by Crossref (Hendricks, Tkaczyk, Lin, & Feeney, 2020). We queried the Crossref REST API6 to collect the DOIs of papers published in volumes 56–83 (February 2018 to June 2021) of this journal.7 We used the Elsevier subscription of the University of Toulouse (GC’s affiliation) to download each article in fulltext XML via the Elsevier API8 and extract the following metadata:

We filtered out publication types other than ‘full-length articles’ and removed two articles with a missing acceptance date. The revision date was missing for 41 articles; we assumed acceptance without revision for these. We noted that no countries were present in the XML format for 12 articles. The final dataset contains 1,078 articles (See Appendix).

We use the term ‘editorial assessment’ to denote the time from submission of a manuscript to its acceptance, including: preliminary screening, invitation of reviewers, rounds of peer review, and final decision. The published metadata for each paper characterises its editorial assessment with three dates: submission, revision, and acceptance. The analysis of the dates of submission vs dates of acceptance reveals a sudden shortening of editorial assessment for volumes published in 2021. Most articles were published after an editorial assessment surprisingly short. Affiliations from China and India were overrepresented. Several blocks of articles shared the same dates of submission and acceptance. These observations depart from the typical publication output of Microprocessors and Microsystems before 2021. Our call for investigation (Sect. 8) invites readers to perform a deeper analysis along the same lines and compare with other reputable journals.

5 Abstracts with high Generative Pre-Training (GPT) detector score 5.1 GPT and the GPT-2 Output Detector The OpenAI company has released several advanced language models: Generative Pretraining (GPT, Radford, Narasimhan, Salimans, & Sutskever, 2018), Generative Pre-trained Transformer 2 (GPT-2, Solaiman, Clark, & Brundage, 2019), and GPT-3 (Brown et al., 2020). The generative power of these models has been extensively discussed:

• “Humans find GPT-2 outputs convincing. Our partners at Cornell University surveyed people to assign GPT-2 text a credibility score across model sizes.” (Solaiman, Clark, & Brundage, 2019) • “We’ve seen no strong evidence of misuse so far.

Discussion. The observed shortening time between submission and acceptance may reflect poor or deficient editorial assessment. Meanwhile, we noted at least two retractions11 in Microprocessors and Microsystems for text duplication, indicating that the journal responds to integrity concerns at least in some cases. The closing of these two retraction notices are similar (only difference: the wording ‘severe abuse’ or ‘misuse’) and read as:

“As such this article represents a (misuse | severe abuse) of the scientific publishing system. The scientific community takes a very strong view on this matter and apologies are offered to readers of the journal that this was not detected during the submission process.”

While the suspected low editorial standards may explain how texts with tortured phrases got published, the process by which those tortured phrases were coined is quite mysterious. It seems improbable, for any skilled scientist, to use a non-standard terminology to refer to well-known concepts in one’s field. In addition, when authors are able to cite the literature that uses the standard terminology, it is unexpected that they switch to a tortured version of the terminology in their own manuscripts.

Our hypothesis is that the observed tortured phrases were coined by misused natural language processing (NLP) tools: automatic translation, automatic re-writing or even automatic generation of text. Today, the vast majority of these tools relies on advanced language models. In the next section we investigate a way to detect the use of such models.

The concentration of abstracts with a high GPT detector score in Microprocessors and Microsystems (Experimental Set) is intriguing. Nonetheless, texts flagged as synthetic by the GPT detector might be scientifically sound. We visually examined several publications from this journal to go beyond automatic screening. The next section reports critical flaws we found in several papers, including nonsensical text featuring tortured phrases, plagiarised text, and image theft. We believe these publications should be considered for retraction as they “represent a severe abuse of the scientific publishing system”, as quoted in the retraction notices reproduced in Sect. 4.4.

6 Critical flaws found in questionable and problematic publications: individual cases All the above quantitative observations suggest that certain editorial processes in several venues were (and might still be) arranged in a non-conventional manner. In order to test this hypothesis, we analysed several individual papers from the journal Microprocessors and Microsystems. For each case presented in this section, we expose various flaws that, in our opinion, are unacceptable in published scientific literature. Our observations include:

• reuse of text and / or images without acknowledgement; • references to non-existing literature; • references to non-existing internal entities of the paper (e.g., theorems and variables in formulas); • sentences for which we failed to infer any meaning.

Excerpts from each case are reported with a score computed by the GPT-2 Output Detector for which “The results start to get reliable after around 50 tokens.” We searched Google Images — either via the ‘Search by image’ feature or by typing in characteristic keywords — for potential earlier occurrences of selected images that appeared most suspect to us (e.g., irrelevant, of poor visual quality) in the papers we inspected. Note that we did not perform this image screening systematically. As of July 8, 2021 there were no citations for the six cases except for Case 5 and Case 6 with one citation each.

Conclusion. and Call for Action We discovered a number of tortured phrases in the scientific literature, mainly in Computer Science. We further studied one specific journal, Microprocessors and Microsystems, affected by the phenomenon. Our study revealed significant, likely questionable, changes in the journal’s operational mode. These changes did not attract much attention, despite being given out by various hints, including:

• abrupt drop of the average / median duration of editorial assessment; • abrupt surge in the number of articles accepted; • acceptance of evidently synthetic texts; • unexpected author affiliations and/or research interests (in authors’ biographies) and/or research background outside of the scope of the venue (e.g., ‘school of musicology’ in a venue on microprocessors).

We specifically discussed the issues with 7 cases covering 8 papers that we are also reporting on PubPeer (Barbour & Stell, 2020). As of 17 June 2021, none of the 1,078 papers had been commented on PubPeer (See Appendix), suggesting that the issues we found went unnoticed. No systematic screening of the papers containing tortured phrases has been performed to date. Nevertheless, we estimate Microprocessors and Microsystems accepted around 500 questionable articles: 389 papers with short duration of editorial assessment in volumes 80– 83, plus additional papers not yet included in a volume. As of June 25, 2021 there were 225 such articles queued ‘in press’ which are ‘accepted, peer reviewed articles that are not yet assigned to volumes/issues, but are citable using DOI.’ Our study revealed that multiple other venues also published papers with tortured phrases (Fig. 1) and abstracts with high GPT detector scores (Tab. 6). Tailoring the fingerprint–query approach used in (Cabanac & Labbé, in press) is a promising way to comb the literature for tortured phrases.17 Preliminary probes show that several thousands of papers with tortured phrases are indexed in major databases. While we managed to identify and retro-engineer several tortured phrases in Computer Science, other tortured phrases related to the concepts of other scientific fields are yet to be exposed.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do hallucinated citations emerge in AI scholarly output? Can AI systems perform peer review as effectively as humans? Can humans reliably detect and resist AI-generated misinformation? Why does polished AI output gain credibility despite fundamental verifiability problems? Can artificial systems establish authority in domains requiring expert judgment? How do writers navigate authorship and delegation with AI? What human oversight must AI research systems have? How can evaluations be made robust against model reward hacking?