Bibliographic record
Abstract
Sir:FigureAlthough we appreciate Dr. Hammond's commentary1 on our article,2 we would like to take the opportunity to respond to and clarify four of the key issues raised. “Overall, the article is highly technical and uses detailed statistical and study design terminology that is difficult to fully understand.” Over the past decade, our team has conducted and published primary research, reviews, and education pieces in the area of patient-reported outcome instruments. Research related to the BREAST-Q is an example of our work.2–5 Our motivation is simple: we feel it is essential that plastic surgeons play a central role in the development and application of patient-reported outcome instruments, especially at a time when interpretation of such data is becoming so vital to quality care.6 It is important that practicing surgeons be exposed to the science of psychometric research, and with this in mind, we respectfully make no apologies for the technical nature of our article. Questionnaire development and validation is inherently complex work. To develop and validate high-quality patient-reported outcome instruments, robust data from large heterogeneous patient cohorts are analyzed using state-of-the-art psychometric methods. Such methods make it possible to distinguish good items from bad and to ensure that the final scales provide reliable, valid, and responsive measurement. Like the iPhone, the BREAST-Q may be simple for people to use, but the underlying design is necessarily intricate and technical. Although our team strove to make the methods in our article as easy to understand and transparent as possible, it behooves the plastic surgery community to become knowledgeable about these research methods. Just as plastic surgeons learn new and complex surgical techniques, they should also be prepared to learn new techniques and terminology in clinical research. As a potentially useful starting point, we would refer Dr. Hammond to our article entitled “The Science behind Quality-of-Life Measurement: A Primer for Plastic Surgeons.”7 “It is possible that one unfortunate byproduct of using the BREAST-Q may well be to actually stifle scientific inquiry.” Helmholtz's famous dictum “all science is measurement” was ably countered by Kelvin's “all science is measurement, but not all measurement is science.” This is no more true than for the human sciences8 and especially in health measurement.9 However, in the same way as Helmholtz was committed to the creation of and use of high-quality experimental data, we (as patient-reported outcome instrument developers) constantly strive to exceed the highest scientific standards to drive the quality of heath measurement in plastic surgery. We hope that our growing body of work in plastic surgery will provide an ever improving evidence base for rigorous patient-reported outcome data collection. Therefore, given the intent and the rigor of research, we find it difficult to imagine a scenario in which the BREAST-Q might actually stifle scientific inquiry. Although scientific inquiry begins with creative ideas and questions, the ultimate aim should be to move researchable ideas into rigorously designed studies. Our team took an idea and, over 5 years of research, developed a patient-reported outcome instrument, which is now available for use by anyone in the plastic surgery community. The BREAST-Q is one of an increasing number of patient-reported outcome tools available to facilitate (not stifle) scientific inquiry in the specialty of plastic surgery. As the measurement of patient-reported outcomes has become an integral component of clinical research and quality improvement efforts in most other specialties, we would encourage the plastic surgery research community to consider including patient-reported outcomes in new studies being designed. However, not all patient-reported outcome instruments are created equal and thus, as we note above, it is essential that plastic surgeons be able to discern what actually makes a quality metric. “It is not unreasonable that a researcher with a specific bias could manipulate the application of the instrument in a manner such that a particular bias is supported.” Bias is an inherent risk in the design of any study, and it is unfortunate that some researchers may manipulate study results. The BREAST-Q is a scientifically credible and clinically meaningful tool designed to help minimize bias. Just like any measurement tool, however, the BREAST-Q will not be able to redeem a poorly designed or badly conducted study. “It remains unclear what the financial implications of the copyright are as it pertains to scientific inquiry.” The BREAST-Q is copyrighted to protect it from modifications by individual users. Any changes to the items or scales would affect the measurement properties of the scales, interfere with accurate raw data scoring, compromise the quality of studies performed using the BREAST-Q, and limit comparability between studies. There is no royalty fee for academic researchers or clinicians who wish to use the BREAST-Q. Andrea L. Pusic, M.D., M.H.S. Memorial Sloan-Kettering Cancer Center, New York, N.Y. Anne F. Klassen, D.Phil. McMaster University, Hamilton, Ontario, Canada Stefan J. Cano, Ph.D. Peninsula College of Medicine and Dentistry, Plymouth, United Kingdom DISCLOSURE Dr. Pusic is a codeveloper of the BREAST-Q, a patient-reported outcome measure owned by Memorial Sloan-Kettering Cancer Center and the University of British Columbia. Based on the inventor-sharing policies of these institutions, Dr. Pusic receives of portion of royalty generated by the use of the measure in industry-sponsored clinical trials.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.020 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.003 | 0.001 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.006 | 0.007 |
| Insufficient payload (model declined to judge) | 0.212 | 0.084 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".