Evaluating cancer patient–reported outcome measures: Readability and implications for clinical use
Bibliographic record
Abstract
BACKGROUND: The benefits of patient-reported outcome measures (PROMs) are well known; however, their readability has come into question because multiple PROMs have been found to be incomprehensible to patients. This is a critical safety and equity consideration because PROMs are increasingly being integrated into routine clinical practice. A key strategy for promoting patient comprehension is the use of plain language. The aim of this study was to determine whether PROMs routinely used in the cancer setting meet plain-language best practices. METHODS: To report the plain-language level of each PROM, readability (Fry Readability Graph, Simple Measure of Gobbledygook, Flesch Reading Ease, and FORCAST) and understandability assessments (Patient Education Materials Assessment Tool [PEMAT] for Printable Materials) were performed. PROMs at grade level 6 or lower and with PEMAT scores greater than 80% were considered to meet plain-language best practices. PROMs were divided into 4 domains (physical, emotional, social, and quality of life) and 17 dimensions (eg, pain was a dimension of the physical domain). A subanalysis was conducted to determine whether specific domains and dimensions were more likely to adhere to plain-language best practices. RESULTS: More than half of the 45 PROMs evaluated (n = 33 [73%]) had a grade level higher than 6. Understandability scores ranged from 29% to 100%. The majority of the PROMs that did not meet plain-language best practices were within the physical and emotional domains and focused on the patient's symptom experience. CONCLUSIONS: This evaluation shows that more than half of the most commonly used cancer PROMs do not meet plain-language best practices. Practice implications include the necessity for plain-language assessment during the PROM validation process, the consideration of plain language in PROM selection, and plain-language review and editing of low-scoring PROMs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".