MétaCan
Menu
Retour à la cohorte
Enregistrement W2397317031 · doi:10.1097/prs.0b013e31828129f4

Reply

2013· letter· en· W2397317031 sur OpenAlexaffabout
Stefan Cano, Anne F. Klassen, Amie Scott, Peter G. Cordeiro, Andrea L. Pusic

Notice bibliographique

RevuePlastic & Reconstructive Surgery · 2013
Typeletter
Langueen
DomaineMedicine
ThématiqueBreast Implant and Reconstruction
Établissements canadiensMcMaster University
Organismes subventionnairesnon disponible
Mots-clésRasch modelItem response theoryClassical test theoryPolytomous Rasch modelRating scaleScale (ratio)Test (biology)PsychologyLevel of measurementConstruct (python library)EconometricsComputer sciencePsychometricsCognitive psychologyStatisticsClinical psychologyMathematicsDevelopmental psychology

Résumé

récupéré en direct d'OpenAlex

Sir:FigureWe would like to thank Dr. Sunil Otiv for his letter1 in response to our article “The BREAST-Q: Further Validation in Independent Clinical Samples.”2 We would agree that Rasch measurement theory has much to offer clinical outcomes assessment in plastic surgery outcomes research, not only in the development and validation of instruments in patient-reported outcomes, but also for clinician-reported and observer-reported outcomes.3,4 We would also like to thank Dr. Otiv for his explanation of how the Rasch model compares to classical test theory. This elaboration is particularly significant, as direct comparisons of these approaches are sparse,5,6 probably because they use different methods, produce different information, and apply different criteria for success and failure. Unlike classical test theory, the aim of Rasch measurement theory is to determine the extent to which observed rating scale data satisfy the measurement model. When the data do not fit the model, we examine the data carefully to try and explain the misfit. This central tenet distinguishes the Rasch measurement theory diagnostic paradigm from other psychometric methods that are based on statistical modeling.4 In this regard, we feel this approach has much in common with clinical practice. Except, of course, as psychometricians we are concerned with items and response options as opposed to signs and symptoms. To take Dr. Otiv's point a little further, we would propose that Rasch measurement theory offers tangible clinical benefits (Table 1), including (when data fit the Rasch model) (1) the ability to construct linear measurements from ordinal-level data, thereby addressing a major concern of using patient-reported outcome instruments as outcome measures7; (2) providing item estimates that are free from the sample distribution and person estimates that are free from the scale distribution, thus allowing for greater flexibility in situations where different samples or scales are used8; and (3) enabling estimates suitable for individual person analyses rather than only for group comparison studies, which speaks directly to surgeons, as the patient is the central unit of interest.9Table 1: Features of Classical Test Theory and Rasch Measurement TheoryWe would also agree with Dr. Otiv that patient-reported outcome instrument validity is key and requires considerable attention. The current U.S. Food and Drug Administration's scientific requirements for patient-reported outcomes in clinical trials10,11 highlight the importance of establishing validity. In particular, the U.S. Food and Drug Administration emphasizes appropriate conceptual frameworks and definitions as being fundamental. These are best achieved using detailed qualitative assessments, which should include evaluating the extent to which a scale's items represent the construct to be measured; establishing the most appropriate item phrasing, structuring, and context; and ensuring consistency in meaning by cognitive debriefing.10 However, traditionally, new scales are developed through the generation of a large pool of items, followed by grouping the items into potential scales, and then (either statistically or thematically) decisions are made as to what construct each group seems to measure, with the subsequent removal of unwanted or irrelevant items. The limitation of this approach is that the scale content, rather than the construct intended for measurement, defines what the scale measures. This makes interpreting its scores in a clinically meaningful way very difficult.12 In developing the BREAST-Q, we selected a range of qualitative methods, including in-depth patient and clinician interviews, literature review, panel meetings, and cognitive debriefing.13,14 However, in addition to these methods, we also strove to develop explicit descriptions of each BREAST-Q scale, to maximize their utility as clinically interpretable tools. As such, the BREAST-Q was developed “bottom-up” (from a construct definition) rather than “top-down” (from a method of grouping items) to ensure that substantive, clinically grounded hypotheses determined scale content. This involved several rounds of iterative qualitative inquiry using the methods described above to establish clinical validity. This approach provides the optimal foundations to fully understand the measurement performance of each of the new scales.15,16 Using detailed qualitative inquiry together with Rasch Measurement Theory to develop the content of the BREAST-Q means that we have a good understanding of the empirical item order across each scale. Thus, we know which items are associated with each and every possible scale score. For example, we previously used the BREAST-Q Reconstruction: Satisfaction with Breasts scale in a multicenter, cross-sectional study of 672 postmastectomy women. We found that women's satisfaction with their breasts was significantly greater among those who received silicone implants (mean score, 64) compared with those who received saline implants (means score, 57).17 We are able to translate these scores as follows: women in the silicone group scored higher up the scale and therefore typically were satisfied with the “look” and “feel” of their reconstructed breasts, whereas women in the saline group scored toward the middle of the scale, and were satisfied with “size” and “look” of their breasts but not how well they “match” or “feel natural.” The ability to provide qualitative statements for each BREAST-Q scale score begins to make their meaning concrete and thus provides a clear base for clinical interpretation. In relation to Dr. Otiv's four questions about the study,2 we make the following remarks. In his first question, Dr. Otiv highlights the mismatch in the BREAST-Q Augmentation Module: Physical Well-Being scale regarding the classical test theory (Cronbach α) and Rasch measurement theory (Person Separation Index) reliability statistics (0.83 and 0.34, respectively). In fact, the Person Separation Index is sensitive to scale-to-sample mistargeting. In this instance, we interpret this result as reflecting that physical well-being is very high in this surgical group and there is a ceiling effect. As we allude to in the article, there are ways to further build on this finding. For example, one route forward would be to expand the content of this scale to attempt to overcome the issue. Although, clinically speaking, this may be counterintuitive, because low physical morbidity would be expected in this group. Dr. Otiv's second question also relates to the reliability statistic, and he provides some interpretation based on the Winsteps program. However, in our study, we used RUMM 2030,18 which uses the Person Separation Index whose values range from 0 to 1, is analogous to the Cronbach α, and can be handled interpretatively in a similar way, bearing in mind the importance of targeting. In his third question, Dr. Otiv asks about person and item standard errors. We did not report the latter because of space restrictions, but this information is available from the authors on request. In terms of the former person standard errors, these can be generated through the Q-Score package, which is freely available with the BREAST-Q (http://webcore.mskcc.org/breastq/scoreBQ.html). Dr. Otiv's final question relates to Dr. Cano's views relating to the relative benefits of Rasch measurement theory and classical test theory. We hope that our position as stated in this letter clears up that issue. Dr. Cano has also expanded on his views elsewhere.12,19 In short, as a research group, we strongly advocate the use of Rasch measurement theory because of its clear clinical benefits over other psychometric methods. As such, we also support Dr. Otiv's four key areas for future debate and expansion surrounding the use of Rasch measurement theory in the development and validation of rating scales in plastic surgery,1 and we believe journals such as Plastic and Reconstructive Surgery are ideally placed to hold such debates. This is because plastic surgeons are key stakeholders in high-stakes clinical outcomes research. As such, they increasingly rely on rating scales to deliver high-quality, reliable, valid, and interpretable measurement. Stefan J. Cano, Ph.D. Peninsula College of Medicine and Dentistry, Plymouth, United Kingdom Anne F. Klassen, D.Phil. McMaster University, Hamilton, Ontario, Canada Amie Scott, B.Sc. Peter G. Cordeiro, M.D. Andrea L. Pusic, M.D., M.H.S. Memorial Sloan-Kettering Cancer Center, New York, N.Y. DISCLOSURE The BREAST-Q is owned by Memorial Sloan-Kettering Cancer Center and the University of British Columbia. Drs. Cano, Klassen, and Pusic are co-developers of the BREAST-Q and receive a portion of the revenues generated when the BREAST-Q is used in industry-sponsored clinical trials.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,002
score de la tête « metaresearch » (Gemma)0,027
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Éditorial · Signal consensuel: aucune
Score de désaccord entre enseignants0,067
Score d'incertitude au seuil0,000

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0020,027
Méta-épidémiologie (sens strict)0,0010,000
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0010,000
Études des sciences et des technologies0,0020,002
Communication savante0,0030,004
Science ouverte0,0020,002
Intégrité de la recherche0,0110,017
Charge utile insuffisante (le modèle a refusé de juger)0,0670,047

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,022
Tête enseignante GPT0,227
Écart entre enseignants0,205 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreÉditorial

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2013
Routes d'admission2
Résumé présentoui

Explorer davantage

Même revuePlastic & Reconstructive SurgeryMême sujetBreast Implant and ReconstructionTravaux en français237 207