Can patient-reported outcomes be used instead of clinician-reported outcomes and photographs as primary endpoints of late normal tissue effects in breast radiotherapy trials? Results from the IMPORT LOW trial
Bibliographic record
Abstract
BACKGROUND: In an era of low local relapse rates after adjuvant breast radiotherapy, risks of late normal-tissue effects (NTE) need to be balanced against risk of relapse. NTE are assessed using patient-reported outcome measures (PROMs), clinician-reported outcomes (CRO) and photographs. This analysis investigates whether PROMs can be used as primary NTE endpoints in breast radiotherapy trials. METHODS: Analyses were conducted within IMPORT LOW (ISRCTN12852634) at 2 and 5 years. NTE were recorded by CRO, photographs and PROMs. Measures of agreement tested concordance, risk ratios for radiotherapy groups were compared, and influence of baseline characteristics on concordance investigated. RESULTS: In 1095 patients who consented to PROMS and photographs, PROMs were available at 2 and/or 5 years for 976 patients, of whom 909 had CRO and 844 had photographs. Few patients had moderate/marked NTE, irrespective of method used (eg. 19% patients and 9% clinicians reported breast shrinkage at year-5). Patients reported more NTE than assessed from CRO or photographs (p < 0.001 for most NTE). Concordance between assessments was poor on an individual patient level; eg. for year-5 breast shrinkage, % agreement = 48% and weighted kappa = 0.17. Risk ratios comparing radiotherapy schedules were consistent between PROMs and CRO or photographs. CONCLUSIONS: Few patients had moderate/marked NTE irrespective of method used. Patients reported more NTE than CRO and photographs, therefore NTE may be underestimated if PROMs are not used. Despite poor concordance between methods, effect sizes from PROMs were consistent with CRO and photographs, suggesting PROMs can be used as primary NTE endpoints in breast radiotherapy trials.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".