[V] Policies on Artificial Intelligence Among Academic Publishers
Bibliographic record
Abstract
Jeremy Y. Ng,<sup>1,2,3</sup> Daivat Bhavsar,<sup>4</sup> Laura Duffy,<sup>4</sup> Hamin Jo,<sup>4</sup> Cynthia Lokker,<sup>4</sup> R. Brian Haynes,<sup>4,5</sup> Alfonso Iorio,<sup>4,5</sup> Ana Marušić<sup>6</sup> <h4>Objective </h4> This study examined the policies implemented by academic publishers regarding authors’ use of generative artificial intelligence (GenAI) tools, focusing on their regulation, disclosure requirements, and role in ensuring the integrity of scientific publications. By analyzing the prevalence and content of these policies, this study aimed to provide insight into the current landscape and inform future policy development in the rapidly evolving field of artificial intelligence (AI)–assisted research and publication. <h4>Design </h4> A cross-sectional audit was conducted on the publicly available policies of 163 academic publishers listed as members of the International Association of Scientific, Technical, and Medical Publishers. Policies were collected and analyzed between September 1 and December 31, 2023. Publishers without publicly accessible policies specific to GenAI use by authors were excluded. Data extraction and analysis were conducted independently in duplicate, with a third reviewer resolving discrepancies. The key policy components analyzed included authorship accreditation, disclosure requirements, and permissions for tasks such as research methods, content generation, image creation, and proofreading. Descriptive statistics were used to summarize the findings. Our protocol was registered.<sup>1</sup> <h4>Results </h4> Of 163 academic publishers, 56 (34.4%) had publicly available policies guiding GenAI use by authors. None permitted authorship accreditation for AI tools, citing accountability concerns and alignment with ethical guidelines. Nearly all publishers with policies (49 of 56 [87.5%]) mandated disclosure of GenAI use, primarily in the Methods or Acknowledgments section. However, disclosure practices varied, with some publishers providing standardized templates while others left requirements vague. Four publishers completely prohibited GenAI use in manuscript preparation, while others allowed their use for specific tasks. Most (33 of 56 [58.9%]) publishers permitted GenAI for drafting nonmethodological sections (eg, Introductions), while 18 (32.1%) permitted their use in research methods, such as data analysis and organization. Few publishers addressed GenAI use in image generation (14 of 163 [8.6%]) or proofreading (15 of 163 [9.2%]). Only 1 publisher (0.6%) allowed citation of AI as primary sources, while 19 (11.6%) explicitly prohibited such citations. Our study has been published.<sup>2</sup> <h4>Conclusions </h4> This audit highlights the inconsistent development of GenAI policies among academic publishers, with large variability in scope and clarity. While the prohibition of AI authorship and the emphasis on mandatory disclosure are consistent themes, inconsistencies in regulating specific tasks suggest a need for standardized and comprehensive policies. As AI technology and its applications in research evolve, publishers must adapt to safeguard scientific integrity. Given that the AI landscape is fast moving, future research includes updating this audit and comparing and contrasting current policies with those found in this study. Future research should also assess how policies are implemented and enforced by examining samples of published articles, as well as explore how these policies affect editors and reviewers, taking into account potential risks, such as privacy breaches and bias. <h4>References</h4> 1. Bhavsar D, Lokker C, Haynes RB, Iorio A, Marusic A, Ng JY. Academic publisher artificial intelligence chatbot policies for authors: a cross-sectional audit. OSF Registries. Accessed July 11, 2025. doi:10.17605/OSF.IO/937ES <span lang="fr-FR">2. Bhavsar D, Duffy L, Jo H, et al. Policies on artificial intelligence chatbots among academic publishers: a cross-sectional audit. </span><i>Res Integr Peer Rev</i><span lang="fr-FR">. 2025;10(1):1. </span>doi:10.1186/s41073-025-00158-y <sup>1</sup>Institute of General Practice and Interprofessional Care, University Hospital Tübingen, Tübingen, Germany, jeremyyng.phd@gmail.com; <sup>2</sup>Robert Bosch Center for Integrative Medicine and Health, Bosch Health Campus, Stuttgart, Germany; <sup>3</sup>Centre for Journalology, Ottawa Hospital Research Institute, Ottawa, Canada; <sup>4</sup>Department of Health Research Methods, Evidence, and Impact, Faculty of Health Sciences, McMaster University, Hamilton, Ontario, Canada; <sup>5</sup>Department of Medicine, McMaster University, Hamilton, Ontario, Canada; <sup>6</sup>Department of Research in Biomedicine and Health and Center for Evidence-Based Medicine, University of Split School of Medicine, Split, Croatia. <h4>Conflicts of Interest Disclosures </h4> The authors declare no conflicts of interest. Ana Marušić is a member of the Peer Review Congress Advisory Board but was not involved in the review or decision for this abstract.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.009 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.013 | 0.021 |
| Science and technology studies | 0.002 | 0.030 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.012 | 0.003 |
| Research integrity | 0.002 | 0.007 |
| Insufficient payload (model declined to judge) | 0.013 | 0.035 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".