Clinical Interpretation Standards and Quality Assurance for the Multicenter PET/CT Trial Rubidium-ARMI
Bibliographic record
Abstract
UNLABELLED: Rubidium-ARMI ((82)Rb as an Alternative Radiopharmaceutical for Myocardial Imaging) is a multicenter trial to evaluate the accuracy, outcomes, and cost-effectiveness of low-dose (82)Rb perfusion imaging using 3-dimensional (3D) PET/CT technology. Standardized imaging protocols are essential to ensure consistent interpretation. METHODS: Cardiac phantom qualifying scans were obtained at 7 recruiting centers. Low-dose (10 MBq/kg) rest and pharmacologic stress (82)Rb PET scans were obtained in 25 patients at each site. Summed stress scores, summed rest scores, and summed difference scores (SSS, SRS, and SDS [respectively] = SSS-SRS) were evaluated using 17-segment visual interpretation with a discretized color map. All scans were coread at the core lab (University of Ottawa Heart Institute) to assess agreement of scoring, clinical diagnosis, and image quality. Scoring differences greater than 3 underwent a third review to improve consensus. Scoring agreement was evaluated with intraclass correlation coefficient (ICC-r), concordance of clinical interpretation, and image quality using κ coefficient and percentage agreement. Patient (99m)Tc and (201)Tl SPECT scans (n = 25) from 2 centers were analyzed similarly for comparison to (82)Rb. RESULTS: Qualifying scores of SSS = 2, SDS = 2, were achieved uniformly at all imaging sites on 9 different 3D PET/CT scanners. Patient scores showed good agreement between core and recruiting sites: ICC-r = 0.92, 0.77 for SSS, SDS. Eighty-five and eighty-seven percent of SSS and SDS scores, respectively, had site-core differences of 3 or less. After consensus review, scoring agreement improved to ICC-r = 0.97, 0.96 for SSS, SDS (P < 0.05). The agreement of normal versus abnormal (SSS ≥ 4) and nonischemic versus ischemic (SDS ≥ 2) studies was excellent: ICC-r = 0.90 and 0.88. Overall interpretation showed excellent agreement, with a κ = 0.94. Image quality was perceived differently by the site versus core reviewers (90% vs. 76% good or better; P < 0.05). By comparison, scoring agreement of the SPECT scans was ICC-r = 0.82, 0.72 for SSS, SDS. Seventy-six and eighty-eight percent of SSS and SDS scores, respectively, had site-core differences of 3 or less. Consensus review again improved scoring agreement to ICC-r = 0.97, 0.90 for SSS, SDS (P < 0.05). CONCLUSION: (82)Rb myocardial perfusion imaging protocols were implemented with highly repeatable interpretation in centers using 3D PET/CT technology, through an effective standardization and quality assurance program. Site scoring of (82)Rb PET myocardial perfusion imaging scans was found to be in good agreement with core lab standards, suggesting that the data from these centers may be combined for analysis of the rubidium-ARMI endpoints.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.014 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".