MétaCan
Menu
Back to cohort
Record W4402535105 · doi:10.1093/bjro/tzae029

Accuracy of an artificial intelligence-enabled diagnostic assistance device in recognizing normal chest radiographs: a service evaluation

2023· article· en· W4402535105 on OpenAlexaff
Amrita Kumar, Puja Patel, Dennis Robert, Shamie Kumar, Aneesh Khetani, Bhargava Reddy, Anumeha Srivastava

Bibliographic record

VenueBJR|Open · 2023
Typearticle
Languageen
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsHendrix Genetics (Canada)
Fundersnot available
KeywordsRadiographyService (business)MedicineComputer scienceMedical physicsArtificial intelligenceRadiologyBusiness

Abstract

fetched live from OpenAlex

Objectives: Artificial intelligence (AI) enabled devices may be able to optimize radiologists' productivity by identifying normal and abnormal chest X-rays (CXRs) for triaging. In this service evaluation, we investigated the accuracy of one such AI device (qXR). Methods: A randomly sampled subset of general practice and outpatient-referred frontal CXRs from a National Health Service Trust was collected retrospectively from examinations conducted during November 2022 to January 2023. Ground truth was established by consensus between 2 radiologists. The main objective was to estimate negative predictive value (NPV) of AI. Results: A total of 522 CXRs (458 [87.74%] normal CXRs) from 522 patients (median age, 64 years [IQR, 49-77]; 305 [58.43%] female) were analysed. AI predicted 348 CXRs as normal, of which 346 were truly normal (NPV: 99.43% [95% CI, 97.94-99.93]). The sensitivity, specificity, positive predictive value, and area under the ROC curve of AI were found to be 96.88% (95% CI, 89.16-99.62), 75.55% (95% CI, 71.34-79.42), 35.63% (95% CI, 28.53-43.23), and 91.92% (95% CI, 89.38-94.45), respectively. A sensitivity analysis was conducted to estimate NPV by varying assumptions of the prevalence of normal CXRs. The NPV ranged from 88.96% to 99.54% as prevalence increased. Conclusions: The AI device recognized normal CXRs with high NPV and has the potential to increase radiologists' productivity. Advances in knowledge: There is a need for more evidence on the utility of AI-enabled devices in identifying normal CXRs. This work adds to such limited evidence and enables researchers to plan studies to further evaluate the impact of such devices.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.003
metaresearch head score (Gemma)0.005
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Other design · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.934
Threshold uncertainty score0.995

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0030.005
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.003
Science and technology studies0.0000.000
Scholarly communication0.0000.001
Open science0.0010.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.401
GPT teacher head0.497
Teacher spread0.096 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designOther design
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueBJR|OpenSame topicArtificial Intelligence in Healthcare and EducationFrench-language works237,207