Measuring What Matters in Radiology: A Guide to Selecting, Implementing, and Interpreting Patient-Reported Outcome Measures
Bibliographic record
Abstract
Patient-reported outcome measures (PROMs) are standardized, validated instruments that measure how patients feel and function, collected directly from the patient. Traditionally, key metrics in radiology include technical aspects such as image quality, radiation dose, and diagnostic accuracy. However, medical imaging and image-guided therapies shape patient experience in informational, emotional, physical, and logistical domains that are rarely measured. Failing to capture this information is an important gap in radiology research and practice today that needs to be addressed. This review synthesizes the science of PROMs through a radiology lens: what PROMs are; why PROMs are relevant to diagnostic imaging and interventional practice; how to select and interpret PROMs responsibly, with explicit attention to bias, conflicts of interest, and minimal important differences; and how to implement PROMs pragmatically using contemporary digital workflows. This article highlights radiology-specific frameworks for patient-centred outcomes of diagnostic tests, summarizes evidence on how electronic PROM (ePROM) programs can improve patient experience and clinical outcomes, and proposes a practical roadmap for department-level implementation. Throughout, this review aligns recommendations with current methodological and regulatory guidance, draws on Canadian implementation experience, and translates lessons from applied PROM programs in complex clinical services to radiology settings. Implemented thoughtfully, PROMs give radiologists a rigorous, low-burden way to document benefits radiology already provides, strengthen outcome and health-economic analyses, and co-design services around what patients value. Integrating PROMs alongside established technical and diagnostic metrics can extend radiology's value proposition, and make radiology's patient-centred impact visible, measurable, and improvable.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.012 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".