Web-Based Patient Educational Material on Osteosarcoma: Quantitative Assessment of Readability and Understandability
Bibliographic record
Abstract
BACKGROUND: Patients often turn to web-based resources following the diagnosis of osteosarcoma. To be fully understood by average American adults, the American Medical Association (AMA) and National Institutes of Health (NIH) recommend web-based health information to be written at a 6th grade level or lower. Previous analyses of osteosarcoma resources have not measured whether text is written such that readers can process key information (understandability) or identify available actions to take (actionability). The Patient Education Materials Assessment Tool (PEMAT) is a validated measurement of understandability and actionability. OBJECTIVE: The purpose of this study was to evaluate web-based osteosarcoma resources using measures of readability, understandability, and actionability. METHODS: Using the search term "osteosarcoma," two independent Google searches were performed on March 7, 2020 (by AGS), and March 11, 2020 (by TRG). The top 50 results were collected. Websites were included if they were directed at providing patient education on osteosarcoma. Readability was quantified using validated algorithms: Flesh-Kincaid Grade Ease (FKGE), Flesch-Kincaid Grade-Level (FKGL). A higher FKGE score indicates that the material is easier to read. All other readability scores represent the US school grade level. Two independent PEMAT assessments were performed with independent scores assigned for both understandability and actionability. A PEMAT score of 70% or below is considered poorly understandable or poorly actionable. Statistical significance was defined as P≤.05. RESULTS: Two searches yielded 53 unique websites, of which 37 (70%) met the inclusion criteria. The mean FKGE and FKGL scores were 40.8 (SD 13.6) and 12.0 (SD 2.4), respectively. No website scored within the acceptable NIH or AHA recommended reading level. Only 4 (11%) and 1 (3%) website met the acceptable understandability and actionability threshold. Both understandability and actionability were positively correlated with FKGE (ρ=0.55, P<.001; ρ=0.60, P<.001), but were otherwise not significantly associated with other readability scores. There were no associations between readability (P=.15), understandability (P=.20), or actionability (P=.31) scores and Google rank. CONCLUSIONS: Overall, web-based osteosarcoma patient educational materials scored poorly with respect to readability, understandability, and actionability. None of the web-based resources scored at the recommended reading level. Only 4 achieved the appropriate score to be considered understandable by the general public. Authors of patient resources should incorporate PEMAT and readability criteria to improve web-based resources to support patient understanding.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.010 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".