MétaCan
Menu
Back to cohort
Record W2397317031 · doi:10.1097/prs.0b013e31828129f4

Reply

2013· letter· en· W2397317031 on OpenAlexaffabout
Stefan Cano, Anne F. Klassen, Amie Scott, Peter G. Cordeiro, Andrea L. Pusic

Bibliographic record

VenuePlastic & Reconstructive Surgery · 2013
Typeletter
Languageen
FieldMedicine
TopicBreast Implant and Reconstruction
Canadian institutionsMcMaster University
Fundersnot available
KeywordsRasch modelItem response theoryClassical test theoryPolytomous Rasch modelRating scaleScale (ratio)Test (biology)PsychologyLevel of measurementConstruct (python library)EconometricsComputer sciencePsychometricsCognitive psychologyStatisticsClinical psychologyMathematicsDevelopmental psychology

Abstract

fetched live from OpenAlex

Sir:FigureWe would like to thank Dr. Sunil Otiv for his letter1 in response to our article “The BREAST-Q: Further Validation in Independent Clinical Samples.”2 We would agree that Rasch measurement theory has much to offer clinical outcomes assessment in plastic surgery outcomes research, not only in the development and validation of instruments in patient-reported outcomes, but also for clinician-reported and observer-reported outcomes.3,4 We would also like to thank Dr. Otiv for his explanation of how the Rasch model compares to classical test theory. This elaboration is particularly significant, as direct comparisons of these approaches are sparse,5,6 probably because they use different methods, produce different information, and apply different criteria for success and failure. Unlike classical test theory, the aim of Rasch measurement theory is to determine the extent to which observed rating scale data satisfy the measurement model. When the data do not fit the model, we examine the data carefully to try and explain the misfit. This central tenet distinguishes the Rasch measurement theory diagnostic paradigm from other psychometric methods that are based on statistical modeling.4 In this regard, we feel this approach has much in common with clinical practice. Except, of course, as psychometricians we are concerned with items and response options as opposed to signs and symptoms. To take Dr. Otiv's point a little further, we would propose that Rasch measurement theory offers tangible clinical benefits (Table 1), including (when data fit the Rasch model) (1) the ability to construct linear measurements from ordinal-level data, thereby addressing a major concern of using patient-reported outcome instruments as outcome measures7; (2) providing item estimates that are free from the sample distribution and person estimates that are free from the scale distribution, thus allowing for greater flexibility in situations where different samples or scales are used8; and (3) enabling estimates suitable for individual person analyses rather than only for group comparison studies, which speaks directly to surgeons, as the patient is the central unit of interest.9Table 1: Features of Classical Test Theory and Rasch Measurement TheoryWe would also agree with Dr. Otiv that patient-reported outcome instrument validity is key and requires considerable attention. The current U.S. Food and Drug Administration's scientific requirements for patient-reported outcomes in clinical trials10,11 highlight the importance of establishing validity. In particular, the U.S. Food and Drug Administration emphasizes appropriate conceptual frameworks and definitions as being fundamental. These are best achieved using detailed qualitative assessments, which should include evaluating the extent to which a scale's items represent the construct to be measured; establishing the most appropriate item phrasing, structuring, and context; and ensuring consistency in meaning by cognitive debriefing.10 However, traditionally, new scales are developed through the generation of a large pool of items, followed by grouping the items into potential scales, and then (either statistically or thematically) decisions are made as to what construct each group seems to measure, with the subsequent removal of unwanted or irrelevant items. The limitation of this approach is that the scale content, rather than the construct intended for measurement, defines what the scale measures. This makes interpreting its scores in a clinically meaningful way very difficult.12 In developing the BREAST-Q, we selected a range of qualitative methods, including in-depth patient and clinician interviews, literature review, panel meetings, and cognitive debriefing.13,14 However, in addition to these methods, we also strove to develop explicit descriptions of each BREAST-Q scale, to maximize their utility as clinically interpretable tools. As such, the BREAST-Q was developed “bottom-up” (from a construct definition) rather than “top-down” (from a method of grouping items) to ensure that substantive, clinically grounded hypotheses determined scale content. This involved several rounds of iterative qualitative inquiry using the methods described above to establish clinical validity. This approach provides the optimal foundations to fully understand the measurement performance of each of the new scales.15,16 Using detailed qualitative inquiry together with Rasch Measurement Theory to develop the content of the BREAST-Q means that we have a good understanding of the empirical item order across each scale. Thus, we know which items are associated with each and every possible scale score. For example, we previously used the BREAST-Q Reconstruction: Satisfaction with Breasts scale in a multicenter, cross-sectional study of 672 postmastectomy women. We found that women's satisfaction with their breasts was significantly greater among those who received silicone implants (mean score, 64) compared with those who received saline implants (means score, 57).17 We are able to translate these scores as follows: women in the silicone group scored higher up the scale and therefore typically were satisfied with the “look” and “feel” of their reconstructed breasts, whereas women in the saline group scored toward the middle of the scale, and were satisfied with “size” and “look” of their breasts but not how well they “match” or “feel natural.” The ability to provide qualitative statements for each BREAST-Q scale score begins to make their meaning concrete and thus provides a clear base for clinical interpretation. In relation to Dr. Otiv's four questions about the study,2 we make the following remarks. In his first question, Dr. Otiv highlights the mismatch in the BREAST-Q Augmentation Module: Physical Well-Being scale regarding the classical test theory (Cronbach α) and Rasch measurement theory (Person Separation Index) reliability statistics (0.83 and 0.34, respectively). In fact, the Person Separation Index is sensitive to scale-to-sample mistargeting. In this instance, we interpret this result as reflecting that physical well-being is very high in this surgical group and there is a ceiling effect. As we allude to in the article, there are ways to further build on this finding. For example, one route forward would be to expand the content of this scale to attempt to overcome the issue. Although, clinically speaking, this may be counterintuitive, because low physical morbidity would be expected in this group. Dr. Otiv's second question also relates to the reliability statistic, and he provides some interpretation based on the Winsteps program. However, in our study, we used RUMM 2030,18 which uses the Person Separation Index whose values range from 0 to 1, is analogous to the Cronbach α, and can be handled interpretatively in a similar way, bearing in mind the importance of targeting. In his third question, Dr. Otiv asks about person and item standard errors. We did not report the latter because of space restrictions, but this information is available from the authors on request. In terms of the former person standard errors, these can be generated through the Q-Score package, which is freely available with the BREAST-Q (http://webcore.mskcc.org/breastq/scoreBQ.html). Dr. Otiv's final question relates to Dr. Cano's views relating to the relative benefits of Rasch measurement theory and classical test theory. We hope that our position as stated in this letter clears up that issue. Dr. Cano has also expanded on his views elsewhere.12,19 In short, as a research group, we strongly advocate the use of Rasch measurement theory because of its clear clinical benefits over other psychometric methods. As such, we also support Dr. Otiv's four key areas for future debate and expansion surrounding the use of Rasch measurement theory in the development and validation of rating scales in plastic surgery,1 and we believe journals such as Plastic and Reconstructive Surgery are ideally placed to hold such debates. This is because plastic surgeons are key stakeholders in high-stakes clinical outcomes research. As such, they increasingly rely on rating scales to deliver high-quality, reliable, valid, and interpretable measurement. Stefan J. Cano, Ph.D. Peninsula College of Medicine and Dentistry, Plymouth, United Kingdom Anne F. Klassen, D.Phil. McMaster University, Hamilton, Ontario, Canada Amie Scott, B.Sc. Peter G. Cordeiro, M.D. Andrea L. Pusic, M.D., M.H.S. Memorial Sloan-Kettering Cancer Center, New York, N.Y. DISCLOSURE The BREAST-Q is owned by Memorial Sloan-Kettering Cancer Center and the University of British Columbia. Drs. Cano, Klassen, and Pusic are co-developers of the BREAST-Q and receive a portion of the revenues generated when the BREAST-Q is used in industry-sponsored clinical trials.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.027
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Editorial · Consensus signal: none
Teacher disagreement score0.067
Threshold uncertainty score0.000

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.027
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.000
Science and technology studies0.0020.002
Scholarly communication0.0030.004
Open science0.0020.002
Research integrity0.0110.017
Insufficient payload (model declined to judge)0.0670.047

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.022
GPT teacher head0.227
Teacher spread0.205 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreEditorial

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2013
Admission routes2
Has abstractyes

Explore more

Same venuePlastic & Reconstructive SurgerySame topicBreast Implant and ReconstructionFrench-language works237,207