What Do Editors-in-Chief of Medical Journals Think About the Use of Artificial Intelligence Chatbots in the Scholarly Publishing Process? Results From An International Cross-Sectional Survey Across Multiple Publishers
Bibliographic record
Abstract
Objective: This study aimed to examine the attitudes and perceptions of Editors-in-Chief (EiCs) of biomedical journals regarding the integration of artificial intelligence chatbots (AICs) into the scholarly publishing process. While AICs offer opportunities to streamline editorial tasks such as plagiarism detection, language editing, and ethics screening, they also introduce ethical, technical, and operational challenges. Understanding EiC perspectives is critical to shaping guidelines, policies, and training programs that align with the evolving role of AICs in scientific publishing. Design: We conducted a cross-sectional survey of EiCs from biomedical journals published by Springer & BMC (part of Springer Nature), Taylor & Francis, Elsevier, Wiley, and SAGE, which are the five largest academic publishers by journal count. Eligible journals were identified through a combination of automated web scraping of publisher webpages and manual verification. A total of 3381 EiCs were invited via email to participate in an anonymous online survey conducted over five weeks in 2024, which included three follow-up reminders. The survey covered familiarity with AICs, current usage, perceived benefits and challenges, and anticipated future roles. Quantitative data were analyzed using descriptive statistics, while qualitative responses underwent thematic content analysis to identify key themes. Results: Of the 3381 EiCs contacted, 510 responded (15.1% response rate), with 505 eligible participants and a completion rate of 87.0%. Most respondents were familiar with AICs (66.7%, 325/487) but had not used them in editorial workflows (83.7%, 401/479). Perceived benefits included enhanced language and grammar support (70.8%, 308/435) and plagiarism screening (67.3%, 294/437). However, respondents expressed concerns about initial setup and training (83.9%, 360/429), ethical risks (80.6%, 345/428), and technical reliability (75.2%, 322/428). While only 49.6% (240/484) of journals reported having formal AIC policies, 89.5% (419/468) of respondents supported training initiatives to promote ethical and effective usage. Despite limited current adoption, 78.9% (370/469) believed AICs will play an important role in the future of scholarly publishing, and 77.2% (363/470) anticipated their significance in advancing scientific research. Themes identified through thematic analysis of open-ended questions include: “no AI in authorship or peer review” referring to the EiC current journal/publisher policy on AIC use, and “ethical, integrity, and privacy concerns” referring to EiC perceptions of challenges with the use of AICs in the scholarly publishing process. Conclusions: Biomedical journal EiCs recognize AICs’ potential to enhance editorial processes but highlight critical barriers, including ethical dilemmas, resource limitations, and insufficient policies and training. Structured interventions, including targeted training programs and robust ethical guidelines, are essential for addressing these challenges and ensuring responsible and effective integration of AICs into publishing workflows.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.027 | 0.130 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".