Reliability and Validity Testing of the Patient and Observer Scar Assessment Scale in Evaluating Linear Scars after Breast Cancer Surgery
Bibliographic record
Abstract
BACKGROUND: The Patient and Observer Scar Assessment Scale is a promising new method incorporating observer and patient ratings in evaluating burn scars. The authors compared this tool to the Vancouver Scar Scale in a cohort of women with linear scars from breast cancer surgery. METHODS: Twenty women with newly diagnosed breast cancer were prospectively accrued. Thirty-one scars were evaluated. The median time from surgery to scar assessment was 8 weeks (range, 3 to 25 weeks). Observer assessment was performed by three independent raters using the Vancouver scale and the observer component of the new tool. Patient self-assessment was performed using the patient component of the tool. Internal consistency, interobserver reliability, and convergent validity were examined. RESULTS: Internal consistency was acceptable for the Vancouver scale and both components of the new tool (Cronbach's alpha, 0.71, 0.74, and 0.77, respectively). Interobserver reliability was substantial with both the Vancouver scale and the observer tool (average measure intraclass coefficient correlation, 0.78 and 0.60, respectively). The observer tool and Vancouver scale correlated significantly with each other (p < 0.001), but only the observer tool correlated well with patients' ratings (p = 0.04). CONCLUSIONS: In surgical scar assessment, the new Patient and Observer Scar Assessment Scale and Vancouver Scar Scale were both associated with acceptable internal consistency and interobserver reliability. The new tool is more comprehensive and has higher correlation with patients' ratings. These findings support the use of the new tool as a reliable, valid, and comprehensive approach to assess linear surgical scars.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.054 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".