Evaluation of a BERT Natural Language Processing Model for Automating CT and MRI Triage and Protocol Selection
Bibliographic record
Abstract
Purpose: To evaluate the accuracy of a Bidirectional Encoder Representations for Transformers (BERT) Natural Language Processing (NLP) model for automating triage and protocol selection of cross-sectional image requisitions. Methods: A retrospective study was completed using 222 392 CT and MRI studies from a single Canadian university hospital database (January 2018-September 2022). Three hundred unique protocols (116 CT and 184 MRI) were included. A BERT model was trained, validated, and tested using an 80%-10%-10% stratified split. Naive Bayes (NB) and Support Vector Machine (SVM) machine learning models were used as comparators. Models were assessed using F1 score, precision, recall, and area under the receiver operating characteristic curve (AUROC). The BERT model was also assessed for multi-class protocol suggestion and subgroups based on referral location, modality, and imaging section. Results: BERT was superior to SVM for protocol selection (F1 score: BERT-0.901 vs SVM-0.881). However, was not significantly different from SVM for triage prediction (F1 score: BERT-0.844 vs SVM-0.845). Both models outperformed NB for protocol and triage. BERT had superior performance on minority classes compared to SVM and NB. For multiclass prediction, BERT accuracy was up to 0.991 for top-5 protocol suggestion, and 0.981 for top-2 triage suggestion. Emergency department patients had the highest F1 scores for both protocol (0.957) and triage (0.986), compared to inpatients and outpatients. Conclusion: The BERT NLP model demonstrated strong performance in automating the triage and protocol selection of radiology studies, showing potential to enhance radiologist workflows. These findings suggest the feasibility of using advanced NLP models to streamline radiology operations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".