MétaCan
Menu
Back to cohort
Record W4213360671 · doi:10.3389/fcvm.2021.816985

A Systematic Quality Scoring Analysis to Assess Automated Cardiovascular Magnetic Resonance Segmentation Algorithms

2022· article· en· W4213360671 on OpenAlexaff
Elisa Rauseo, Muhammad Omer, Alborz Amir-Khalili, Alireza Sojoudi, Thu‐Thao Le, Stuart A. Cook, Derek J. Hausenloy, Briana Ang, Desiree‐Faye Toh, Jennifer Bryant, José Miguel Paiva, Kenneth Fung, Jackie Cooper, Mohammed Y Khanji, Nay Aung, Steffen E. Petersen

Bibliographic record

VenueFrontiers in Cardiovascular Medicine · 2022
Typearticle
Languageen
FieldMedicine
TopicCardiac Imaging and Diagnostics
Canadian institutionsCircle Cardiovascular Imaging
FundersInnovate UKEngineering and Physical Sciences Research CouncilBritish Heart Foundation
KeywordsMagnetic resonance imagingSegmentationCardiac magnetic resonanceComputer scienceQuality (philosophy)MedicineArtificial intelligenceAlgorithmRadiologyPhysics

Abstract

fetched live from OpenAlex

Background The quantitative measures used to assess the performance of automated methods often do not reflect the clinical acceptability of contouring. A quality-based assessment of automated cardiac magnetic resonance (CMR) segmentation more relevant to clinical practice is therefore needed. Objective We propose a new method for assessing the quality of machine learning (ML) outputs. We evaluate the clinical utility of the proposed method as it is employed to systematically analyse the quality of an automated contouring algorithm. Methods A dataset of short-axis (SAX) cine CMR images from a clinically heterogeneous population (n = 217) were manually contoured by a team of experienced investigators. On the same images we derived automated contours using a ML algorithm. A contour quality scoring application randomly presented manual and automated contours to four blinded clinicians, who were asked to assign a quality score from a predefined rubric. Firstly, we analyzed the distribution of quality scores between the two contouring methods across all clinicians. Secondly, we analyzed the interobserver reliability between the raters. Finally, we examined whether there was a variation in scores based on the type of contour, SAX slice level, and underlying disease. Results The overall distribution of scores between the two methods was significantly different, with automated contours scoring better than the manual (OR (95% CI) = 1.17 (1.07–1.28), p = 0.001; n = 9401). There was substantial scoring agreement between raters for each contouring method independently, albeit it was significantly better for automated segmentation (automated: AC2 = 0.940, 95% CI, 0.937–0.943 vs manual: AC2 = 0.934, 95% CI, 0.931–0.937; p = 0.006). Next, the analysis of quality scores based on different factors was performed. Our approach helped identify trends patterns of lower segmentation quality as observed for left ventricle epicardial and basal contours with both methods. Similarly, significant differences in quality between the two methods were also found in dilated cardiomyopathy and hypertension. Conclusions Our results confirm the ability of our systematic scoring analysis to determine the clinical acceptability of automated contours. This approach focused on the contours' clinical utility could ultimately improve clinicians' confidence in artificial intelligence and its acceptability in the clinical workflow.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.111
metaresearch head score (Gemma)0.226
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Systematic review · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.111
Threshold uncertainty score0.588

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1110.226
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.003
Bibliometrics0.0130.006
Science and technology studies0.0010.002
Scholarly communication0.0030.002
Open science0.0020.004
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.024
GPT teacher head0.301
Teacher spread0.276 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSystematic review
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueFrontiers in Cardiovascular MedicineSame topicCardiac Imaging and DiagnosticsFrench-language works237,207