Observed and expected reliability of echocardiographic volumetric methods and critical change values for quantification of mitral regurgitant fraction in dogs
Bibliographic record
Abstract
Abstract Background Reliability of echocardiographic calculations for stroke volume and mitral regurgitant fraction (RFMR) are affected by observer variability and lack of a gold standard. Variability is used to calculate critical change values (CCVs) that are thresholds representing real change in a measure not associated with observer variability. Hypothesis Observed intra- and interobserver accuracy and variability in healthy dogs help model CCV for RFMR. Animals Reliability cohort of 34 healthy dogs; allometric scaling cohort of 99 dogs with heart disease and 25 healthy dogs. Methods Accuracy, variability, and CCV of 2 observers using geometric and flow-based echocardiography were prospectively compared against a standard of RFMR = 0% and extrapolated across a range of expected RFMR values in the reliability cohort partly derived from cardiac dimensions predicted by the allometric cohort. Results Accuracy of methods to determine RFMR in descending order was 4-chamber bullet (Bullet4CH), mitral inflow, cube formula, and Simpson's method of disks. Intraobserver variability was relatively high. The CCV for RFMR ranged from 28% to 88% and was inversely related to RFMR when extrapolated for use in affected dogs. For both observers, the Bullet4CH method had the lowest intraobserver CCV (Operator 1:28%, Operator 2:41%). Interobserver strength of agreement was low with intraclass correlation coefficients ranging from 0.210 to 0.413. Conclusions and Clinical Importance Echocardiographic volumetric methods used to calculate stroke volume and RFMR have low accuracy and high variability in healthy dogs. Extrapolation of observed CCV to a range of expected RFMR suggests observers and methods are not interchangeable and variability might hinder routine clinical usage. Individual observers should be aware of their own variability and CCV.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".