MétaCan
Menu
Back to cohort
Record W4413408952 · doi:10.1007/s00415-025-13261-3

AI-Based EMG Reporting: A Randomized Controlled Trial

2025· article· en· W4413408952 on OpenAlexaff
Alon Gorenshtein, Yana Weisblat, Mohamed Khateb, Gilad Kenan, Irina Tsirkin, Galina Fayn, Shahar Shelly

Bibliographic record

VenueJournal of Neurology · 2025
Typearticle
Languageen
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsUniversity of TorontoUniversity Health Network
FundersBar-Ilan University
KeywordsRandomized controlled trialNeuroradiologyNeurologyMedicinePhysical medicine and rehabilitationPhysical therapyPsychologySurgeryPsychiatry

Abstract

fetched live from OpenAlex

BACKGROUND AND OBJECTIVES: Accurate interpretation of electrodiagnostic (EDX) studies is essential for the diagnosis and management of neuromuscular disorders. Artificial intelligence (AI) based tools may improve consistency and quality of EDX reporting and reduce workload. The aim of this study is to evaluate the performance of an AI-assisted, multi-agent framework (INSPIRE) in comparison with standard physician interpretation in a randomized controlled trial (RCT). METHODS: We prospectively enrolled 200 patients (out of 363 assessed for eligibility) referred for EDX. Patients were randomly assigned to either a control group (physician-only interpretation) or an intervention group (physician-AI interpretation). Three board-certified physicians, rotated across both arms. In the intervention group, an AI-generated preliminary report was combined with the physician's independent findings (human-AI integration). The primary outcome was EDX report quality using score we developed named AI-Generated EMG Report Score (AIGERS; score range, 0-1, with higher scores indicating more accurate or complete reports). Secondary outcomes included a physician-reported AI integration rating score (PAIR) and a compliance survey evaluating ease of AI adoption. RESULTS: Of the 200 enrolled patients, 100 were allocated to AI-assisted interpretation and 100 to physician-only reporting. While AI-generated preliminary reports offered moderate consistency on the AIGERS metric, the integrated (physician-AI) approach did not significantly outperform physician-only methods. Despite some anecdotal advantages such as efficiency in suggesting standardized terminology quantitatively, the AIGERS scores for physician-AI integration was nearly the same as those in the physician-only arm and did not reach statistical significance (p > 0.05 for all comparisons). Physicians reported variable acceptance of AI suggestions, expressing concerns about the interpretability of AI outputs and workflow interruptions. Physician-AI collaboration scores showed moderate trust in the AI's suggestions (mean 3.7/5) but rated efficiency (2.0/5), ease of use (1.7/5), and workload reduction (1.7/5) as poor, indicating usability challenges and workflow interruptions. DISCUSSION: In this single-center, randomized trial, AI-assisted EDX interpretation did not demonstrate a significant advantage over conventional physician-only interpretation. Nevertheless, the AI framework may help reduce workload and documentation burdens by handling simpler, routine EDX tests freeing physicians to focus on more complex cases that require greater expertise. TRIAL REGISTRATION: ClinicalTrials.gov Identifier: NCT06902675.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.003
metaresearch head score (Gemma)0.019
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Randomized trial · Consensus signal: Randomized trial
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.115
Threshold uncertainty score0.989

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0030.019
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.118
GPT teacher head0.470
Teacher spread0.352 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designRandomized trial
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations13
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJournal of NeurologySame topicArtificial Intelligence in Healthcare and EducationFrench-language works237,207