Whole-body magnetic resonance imaging (WB-MRI) reporting with the METastasis Reporting and Data System for Prostate Cancer (MET-RADS-P): inter-observer agreement between readers of different expertise levels
Bibliographic record
Abstract
BACKGROUND: The METastasis Reporting and Data System for Prostate Cancer (MET-RADS-P) guidelines are designed to enable reproducible assessment in detecting and quantifying metastatic disease response using whole-body magnetic resonance imaging (WB-MRI) in patients with advanced prostate cancer (APC). The purpose of our study was to evaluate the inter-observer agreement of WB-MRI examination reports produced by readers of different expertise when using the MET-RADS-P guidelines. METHODS: Fifty consecutive paired WB-MRI examinations, performed from December 2016 to February 2018 on 31 patients, were retrospectively examined to compare reports by a Senior Radiologist (9 years of experience in WB-MRI) and Resident Radiologist (after a 6-months training) using MET-RADS-P guidelines, for detection and for primary/dominant and secondary response assessment categories (RAC) scores assigned to metastatic disease in 14 body regions. Inter-observer agreement regarding RAC score was evaluated for each region by using weighted-Cohen's Kappa statistics (K). RESULTS: The number of metastatic regions reported by the Senior Radiologist (249) and Resident Radiologist (251) was comparable. For the primary/dominant RAC pattern, the agreement between readers was excellent for the metastatic findings in cervical, dorsal, and lumbosacral spine, pelvis, limbs, lungs and other sites (K:0.81-1.0), substantial for thorax, retroperitoneal nodes, other nodes and liver (K:0.61-0.80), moderate for pelvic nodes (K:0.56), fair for primary soft tissue and not assessable for skull due to the absence of findings. For the secondary RAC pattern, agreement between readers was excellent for the metastatic findings in cervical spine (K:0.93) and retroperitoneal nodes (K:0.89), substantial for those in dorsal spine, pelvis, thorax, limbs and pelvic nodes (K:0.61-0.80), and moderate for lumbosacral spine (K:0.44). CONCLUSIONS: We found inter-observer agreement between two readers of different expertise levels to be excellent in bone, but mixed in other body regions. Considering the importance of bone metastases in patients with APC, our results favor the use of MET-RADS-P in response to the growing clinical need for monitoring of metastasis in these patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.042 | 0.062 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".