MétaCan
Menu
Back to cohort
Record W7117672484 · doi:10.1097/scs.0000000000012366

Detectable AI-Generated Text in the Journal of Craniofacial Surgery: A Comparison of 2014 and 2024 Publications

2025· article· en· W7117672484 on OpenAlexaff
Forrest Bohler, Tyler Tran, James R Burmeister, Karam Hadid, Spruha Joshi, Samuel Gao, Kongkrit Chaiyasate

Bibliographic record

VenueJournal of Craniofacial Surgery · 2025
Typearticle
Languageen
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsMcMaster University
Fundersnot available
KeywordsInterpretation (philosophy)CraniofacialGenerative grammarMEDLINESoftware

Abstract

fetched live from OpenAlex

BACKGROUND: The rapid emergence of large language models (LLMs) has transformed scientific writing, prompting concerns regarding the extent to which generative artificial intelligence (AI) tools may be influencing published research manuscripts. AI-detection software has been proposed as a method to identify AI-generated text; however, the validity of these tools in scientific contexts remains uncertain. This study evaluates detectable AI content in articles published in the Journal of Craniofacial Surgery (JCFS) before and after the widespread adoption of LLMs. METHODS: A retrospective cross-sectional analysis was conducted using all JCFS articles published in 2014 (pre-LLM) and 2024 (post-LLM). Full-text manuscripts and individual sections (Abstract, Introduction, Methods, Results, Discussion, Conclusion) were analyzed using ZeroGPT to determine the percentage of detectable AI-generated text. Detection scores were compared using the Mann-Whitney U test. RESULTS: A total of 1490 manuscripts were analyzed (659 in 2014; 831 in 2024). Mean detectable AI content increased from 8.6% (SD 9.8) in 2014 to 10.7% (SD 10.4) in 2024 ( P = 0.00001). Section-level comparison demonstrated the greatest increase in Results sections (19.8%-24.1%, P = 0.00001), with additional increases in Introduction, Methods, Discussion, and Conclusion sections, but no significant change in Abstracts (13.4% versus13.9%, P = 0.32). Although statistically significant, these differences were small in absolute magnitude. CONCLUSIONS: Detectable AI content in JCFS manuscripts increased modestly over the past decade, likely reflecting detection software behavior and evolving writing structure rather than widespread use of generative AI. Findings support cautious interpretation of AI-detection outputs and highlight the need for validated tools and thoughtful editorial policy development.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.022
metaresearch head score (Gemma)0.151
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Research integrity
Consensus categoriesnone
DomainCandidate signal: Evaluation · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.999
Threshold uncertainty score0.119

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0220.151
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0270.018
Science and technology studies0.0010.002
Scholarly communication0.0040.003
Open science0.0010.003
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0040.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.114
GPT teacher head0.419
Teacher spread0.305 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designObservational
DomainEvaluation
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Craniofacial SurgerySame topicArtificial Intelligence in Healthcare and EducationFrench-language works237,207