MétaCan
Menu
Back to cohort
Record W3123388938 · doi:10.1186/s40635-021-00380-0

Evaluation of automated microvascular flow analysis software AVA 4: a validation study

2021· article· en· W3123388938 on OpenAlexaff
Christian S. Guay, Mariam Khebir, Shiva Shahiri, Ariana Szilágyi, Erin Elizabeth Cole, Gabrielle Simoneau, Mohamed Badawy

Bibliographic record

VenueIntensive Care Medicine Experimental · 2021
Typearticle
Languageen
FieldMedicine
TopicHemodynamic Monitoring and Therapy
Canadian institutionsMcGill UniversityMontreal Neurological Institute and Hospital
Fundersnot available
KeywordsMedicineIntraclass correlationReproducibilityGold standard (test)Bland–Altman plotPerioperativeLimits of agreementIntensive careMedical physicsNuclear medicineRadiologyStatisticsIntensive care medicineMathematics

Abstract

fetched live from OpenAlex

BACKGROUND: Real-time automated analysis of videos of the microvasculature is an essential step in the development of research protocols and clinical algorithms that incorporate point-of-care microvascular analysis. In response to the call for validation studies of available automated analysis software by the European Society of Intensive Care Medicine, and building on a previous validation study in sheep, we report the first human validation study of AVA 4. METHODS: Two retrospective perioperative datasets of human microcirculation videos (P1 and P2) and one prospective healthy volunteer dataset (V1) were used in this validation study. Video quality was assessed using the modified Microcirculation Image Quality Selection (MIQS) score. Videos were initially analyzed with (1) AVA software 3.2 by two experienced investigators using the gold standard semi-automated method, followed by an analysis with (2) AVA automated software 4.1. Microvascular variables measured were perfused vessel density (PVD), total vessel density (TVD), and proportion of perfused vessels (PPV). Bland-Altman analysis and intraclass correlation coefficients (ICC) were used to measure agreement between the two methods. Each method's ability to discriminate between microcirculatory states before and after induction of general anesthesia was assessed using paired t-tests. RESULTS: Fifty-two videos from P1, 128 videos from P2 and 26 videos from V1 met inclusion criteria for analysis. Correlational analysis and Bland-Altman analysis revealed poor agreement and no correlation between AVA 4.1 and AVA 3.2. Following the induction of general anesthesia, TVD and PVD measured using AVA 3.2 increased significantly for P1 (p < 0.05) and P2 (p < 0.05). However, these changes could not be replicated with the data generated by AVA 4.1. CONCLUSIONS: AVA 4.1 is not a suitable tool for research or clinical purposes at this time. Future validation studies of automated microvascular flow analysis software should aim to measure the new software's agreement with the gold standard, its ability to discriminate between clinical states and the quality thresholds at which its performance becomes unacceptable.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.002
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: Bench or experimental
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.171
Threshold uncertainty score0.696

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.002
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.038
GPT teacher head0.386
Teacher spread0.348 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations17
Published2021
Admission routes1
Has abstractyes

Explore more

Same venueIntensive Care Medicine ExperimentalSame topicHemodynamic Monitoring and TherapyFrench-language works237,207