Patient-Reported Outcome Measurement Compared with Professional Judgment of Cosmetic Results after Breast-Conserving Therapy
Bibliographic record
Abstract
Background: In the present study, we set out to compare patient-reported outcomes with professional judgment about cosmesis after breast-conserving therapy (bct) and to evaluate which items (position of the nipple, color, scar, size, shape, and firmness) correlate best with subjective outcome. Methods: Dutch patients treated with bct between 2008 and 2009 were analyzed. Exclusion criteria were prior amputation or bct of the contralateral breast, metastatic disease, local recurrence, or any prior cosmetic breast surgery. Structured questionnaires and standardized six-view photographs were obtained with a minimum of 3 years' follow-up. Cosmetic outcome was judged by the patients and, based on photographs, by 5 different medical professionals using 3 different scoring systems: the Harvard scale, the Sneeuw questionnaire, and a numeric rating scale. Agreement was scored using the intraclass correlation coefficient (icc). The association between items of the Sneeuw questionnaire and a fair-poor Harvard score was estimated using logistic regression analysis. Results: The study included 108 female patients (age: 40-91 years). Based on the Harvard scale, agreement on cosmetic outcome between the professionals was good (icc: 0.78). In contrast, agreement between professionals as a group compared with the patients was found to be fair to moderate (icc range: 0.38-0.50). The items "size" and "shape" were identified as the strongest determinants of cosmetic outcome. Conclusions: Cosmetic outcome was scored differently by patients and professionals. Agreement was greater between the professionals than between the patients and the professionals as a group. In general, size and shape were the most prominent items on which cosmetic outcome was judged by patients and professionals alike.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".