Patient-Reported Outcome Measurement Compared with Professional Judgment of Cosmetic Results after Breast-Conserving Therapy
Bibliographic record
Abstract
Background: In the present study, we set out to compare patient-reported outcomes with professional judgment about cosmesis after breast-conserving therapy (bct) and to evaluate which items (position of the nipple, color, scar, size, shape, and firmness) correlate best with subjective outcome. Methods: Dutch patients treated with bct between 2008 and 2009 were analyzed. Exclusion criteria were prior amputation or bct of the contralateral breast, metastatic disease, local recurrence, or any prior cosmetic breast surgery. Structured questionnaires and standardized six-view photographs were obtained with a minimum of 3 years' follow-up. Cosmetic outcome was judged by the patients and, based on photographs, by 5 different medical professionals using 3 different scoring systems: the Harvard scale, the Sneeuw questionnaire, and a numeric rating scale. Agreement was scored using the intraclass correlation coefficient (icc). The association between items of the Sneeuw questionnaire and a fair-poor Harvard score was estimated using logistic regression analysis. Results: The study included 108 female patients (age: 40-91 years). Based on the Harvard scale, agreement on cosmetic outcome between the professionals was good (icc: 0.78). In contrast, agreement between professionals as a group compared with the patients was found to be fair to moderate (icc range: 0.38-0.50). The items "size" and "shape" were identified as the strongest determinants of cosmetic outcome. Conclusions: Cosmetic outcome was scored differently by patients and professionals. Agreement was greater between the professionals than between the patients and the professionals as a group. In general, size and shape were the most prominent items on which cosmetic outcome was judged by patients and professionals alike.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.018 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".