Detectable AI-Generated Text in the Journal of Craniofacial Surgery: A Comparison of 2014 and 2024 Publications
Bibliographic record
Abstract
BACKGROUND: The rapid emergence of large language models (LLMs) has transformed scientific writing, prompting concerns regarding the extent to which generative artificial intelligence (AI) tools may be influencing published research manuscripts. AI-detection software has been proposed as a method to identify AI-generated text; however, the validity of these tools in scientific contexts remains uncertain. This study evaluates detectable AI content in articles published in the Journal of Craniofacial Surgery (JCFS) before and after the widespread adoption of LLMs. METHODS: A retrospective cross-sectional analysis was conducted using all JCFS articles published in 2014 (pre-LLM) and 2024 (post-LLM). Full-text manuscripts and individual sections (Abstract, Introduction, Methods, Results, Discussion, Conclusion) were analyzed using ZeroGPT to determine the percentage of detectable AI-generated text. Detection scores were compared using the Mann-Whitney U test. RESULTS: A total of 1490 manuscripts were analyzed (659 in 2014; 831 in 2024). Mean detectable AI content increased from 8.6% (SD 9.8) in 2014 to 10.7% (SD 10.4) in 2024 ( P = 0.00001). Section-level comparison demonstrated the greatest increase in Results sections (19.8%-24.1%, P = 0.00001), with additional increases in Introduction, Methods, Discussion, and Conclusion sections, but no significant change in Abstracts (13.4% versus13.9%, P = 0.32). Although statistically significant, these differences were small in absolute magnitude. CONCLUSIONS: Detectable AI content in JCFS manuscripts increased modestly over the past decade, likely reflecting detection software behavior and evolving writing structure rather than widespread use of generative AI. Findings support cautious interpretation of AI-detection outputs and highlight the need for validated tools and thoughtful editorial policy development.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.151 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.027 | 0.018 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".