Writing for publication—Implications of text recycling and cut and paste writing
Bibliographic record
Abstract
Increased use and availability of electronic text-matching software (such as ithenticate™) to screen submissions to peer-reviewed journals mean that Editors are increasingly being made aware of potential cases of plagiarism (Debnath, 2016). While there are no true figures for academic plagiarism within submissions to peer-reviewed journals (Gasparyan et al., 2017), anecdotally upwards of 25% of all submissions require some level of follow-up and investigation. The text matching details provided by software such as ithenticate™ support Editor decision-making, and assist Editors to explore all incidences of text matching fully. This can be an onerous and time-consuming task as Editors are finding large numbers of submissions require close attention or further investigation (Hong, 2017). In some cases, Editors have applied percentage cut-off points (i.e. applying a minimal threshold), and rejecting papers on this basis; however, this can be found to be overly punitive or conversely situations, where academic integrity was potentially compromised, may be missed. While approaches to this vary, Editors are increasingly closely examining all submitted papers regardless of the percentage level of matching. An ithenticate™ report is available on all articles submitted to the Journal of Nursing Management (2019) for example, and submissions are examined in this context. However, Masic (2012:212) suggests that there is “a dilemma: who, on what basis (criteria, standards, rules), when and how should declare someone a plagiarist.” As Journal Editors are usually familiar with COPE (2019a,2019c,2019b) guidelines and experienced with dealing with academic misconduct, this dilemma arises more so in cases that are not straightforward duplication or academic dishonesty, but rather the finding of a mosaic match that is usually due to cut and paste writing with or without self-plagiarism (Timmins, 2019). Occasionally, there are matches to the authors’ theses or conference presentations, usually not disclosed to the Editor. In all cases, there are implications for Editors when regularly using text-matching software (Timmins, 2019). A proliferation and increase of publications about the subject of plagiarism in academic journals in recent years (Debnath & Cariappa, 2018) reflect a growing concern about this from both author and Editor perspective. Editorials by Editors and letters from authors are particularly commonplace (Debnath & Cariappa, 2018). Overall, the management of cases can be challenging and the process for navigation is not always clear. The Committee on Publication Ethics (COPE 2019a) provide a useful to Editors that guide and support decision-making in these situations. However, while these ethical guidelines exist (COPE, 2019a,2019c,2019b), few journals provide specific information on plagiarism (Debnath & Cariappa, 2018) and thus authors are not always familiar with Editorial gatekeeping practices or the consequences. There are also calls for a more “uniform international policy” on this across journals (Arumugam & Aldhafiri, 2016:428). Interestingly, the International Committee of Medical Journal Editors (2016) have developed their own specific guidance that is used alongside COPE guidelines. However, while blatant cases of duplication of papers or academic misconduct are reasonably straightforward to detect and manage, supported by the COPE guidelines, cases of blatant duplication (whether self-plagiarism or not) are less common. More commonly, a match is detected with the authors’ theses, presentations or other publications by the author. Usually, the authors provide no information to Editors about potential content overlap although many submission systems now prompt the author to provide such information thus prompting the Editor. If this information is not available or forthcoming, it is incumbent on the Editor then to follow-up individual matches to inform their decision as to the nature and extent of the match, whether or not it constitutes plagiarism (COPE, 2019a, 2019c, 2019b, b, c). Often the Editor finds an undeclared match with the author's thesis or conference presentation prompting Editors to suggest that authors cite the original source (if relevant to the paper). More commonly, the Editor finds cut and paste writing from either the author's own previously published work or from one or more published papers by other authors. This can appear as text recycling (using the author's previously published material to develop a discussion on the new submission) or salami slicing (where findings or methods from a single study are being reworked to produce a novel paper with new or expanded findings and thus redundant) (COPE, 2019c; Rohwer, Young, Wager, & Garner, 2017). In either case, the authors are usually embarking on “copy-and-paste writing” (Gasparyan et al., 2017:1,220) from their own work leading to cross matches with their publications. These matches can range from a small array of sentences and paragraphs that are inadequately paraphrased to discussions/conclusions that are not only paraphrased inappropriately but rely heavily conceptually on the ideas of the author's previously published work (COPE, 2019c). Among the various forms of academic misconduct, self-plagiarism is particularly problematic (Horback & Halffman, 2019). Self-plagiarism is also known as “academic text recycling” and it involves the “reuse of one's own writing in academic publications, ranging from a sentence to several pages” (Horback & Halffman, 2019:492). Recently, Horbach and Halffman (2019) found rates of self-plagiarism in published papers to be at least 6% and up to 40% among some disciplines. Rohwer et al. (2017:1) found that at least 60% of their participating (n = 583) corresponding authors from low/middle-income countries admitted to “text recycling” and found nothing wrong with this. These authors report various issues arising with this including poor paraphrasing; self-plagiarism of online abstracts, reports and theses; self-plagiarism from other author publications; and salami slicing and duplication. Some authors object to the term self-plagiarism, “as stealing from oneself is a legal oxymoron“ (Horback & Halffman, 2019:492), without realising that these words may no longer their own as they may have transferred copyright to the publisher. Although in some cases, small amounts of material overlap are unavoidable and/or acceptable (COPE, 2019c). Again issues arise not in the simultaneous publication of the same paper (as this is relatively straightforward to manage) but rather with the reuse of sections of previously published work (text recycling as detected by text-matching software). Directly copying of the conclusion is of particular consequence as Editors may question what is novel about the paper and what this new conclusion possibly adds to the body of knowledge (even if the remainder of the paper is not copied) (COPE, 2019c). An example where this might occur is where an author previously published a concept analysis on the topic and then proceeds to submit a research study with the inclusion of large portions of the aforementioned publication (Smith, 2016). Self-plagiarism is particularly problematic for Editors, where authors have made no disclosure of simultaneous or overlapping papers (COPE, 2019b). It beholds the Editor then to explore the text-matching report and each individual item and try to interpret whether or not academic integrity is compromised (COPE, 2019a,2019c,2019b). Recent COPE (2019b) guidelines provide guidance on action with regard to text recycling, which, if deemed appropriate by the Editor may be dealt with by simple rewording of text, citing of relevant corresponding papers and resubmission by the authors. Concerns arising with copy and paste writing, whether self-plagiarism or not, relate not only to academic integrity and ethics, but the fact that poor writing skills are being used (Timmins, 2019). The notion of “copy-and-paste writing” (Gasparyan et al., 2017:1,220) or “rogeting” [substituting words when using quotations] is both widely inappropriately understood as academic writing. Many authors, even seasoned academics, appear unused to and using their own voice and “language” to interpret and explain (Debnath, 2016:166) or change the language from that previously used. This point is interesting as although it is commonly postulated that plagiarism is more common among more inexperienced colleagues, objective evidence finds that the opposite is true (Horbach & Halffman, 2019). Quite aside from the issues of plagiarism and possible poor writing technique, when authors have cut and paste from other articles/books or use this “rogeting” approach, it is likely that the emerging discussion is (a) out of date/not current and (b) not novel, even when they are reusing their own material. Thus, a major issue with not interpreting scholarly work in an authentic and novel way is that the paper fails to achieve novelty or currency, two of the main requirements for success at peer-reviewed publication. Failure to achieve these targets can occur with even the most minimal (percentage) of “cut and paste” writing. If, for example, the background/discussion or conclusion of the newly submitted paper is largely a reworking of a previously published concept analysis (Smith, 2016), one is led to question what is new about this material (COPE, 2019c). Additionally, as the original paper is already published, there is likely that a lot of new material is not captured given the passage of time. Overall, the management of the cases is challenging and the process for navigation is not always clear. The effect on and response of authors vary, and place Editors in an invidious situation (Hausmann & Murphy, 2016). Raising awareness about the inappropriateness of this cut and paste writing is challenging with some authors adopting a defensive stance, understandably as it may appear that academic integrity is being questioned. Defences such as “I didn't plagiarise - we used commonly used terms,” “I included a citation/reference,” “I cant express these points in any other way” “this is my concept analysis and I can't change it” etc are common. often authors ask what the journal threshold (percentage accepted plagiarism) is and expect that the same allowances that are made for undergraduate students in their organisations be applied (30% for example), without understanding the Editor's role in applying the COPE (2019a,2019c,2019b) guidelines. However, the reason that following along with usual guidelines for university-based students and applying a cut-off percentage is not appropriate for scientific publications is that authors are expected to comply with the highest level of integrity and ethical standards (Taylor, 2017) and role model excellence in academic writing. Additionally, most journals are striving for novelty, and text overlap severely compromises this in most cases. However, awareness of the issue is growing and the UK Cancer Institute for example now employs personnel to manage data overlap and ensures absolute integrity (Winchester, 2018). Increased use of electronic text-matching software means that Editors are increasingly being challenged to manage potential plagiarism situations effectively. Up to a quarter of submissions require some level of follow-up. Text recycling is common and more attention is being paid to this in recent years with the advent of electronic monitoring of text matching and COPE (2019c) provide guidance on managing this. Of growing concern is the practice of “copy-and-paste writing” (Gasparyan et al., 2017:1,220), perhaps invisible in the past, but increasingly obvious with increased use of text-matching software. Cut and paste writing needs to be avoided, for reasons of integrity but also to ensure novelty in application, approach and conclusions. To this end, researchers need to be encouraged to use their own voice and “language” to interpret and explain materials and concepts (Debnath, 2016:166). Overall, the management of the cases is challenging and the process for navigation is not always clear. This paper discusses some arising issues, outlines current guidance and opens a platform for debate and discussion among researchers, authors and Editors within nursing journals internationally.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".