MétaCan
Menu
Back to cohort
Record W2776532933 · doi:10.1093/asj/sjx063

To Get the Best Outcome, Choose the Best Outcome

2017· letter· en· W2776532933 on OpenAlexaff
Achilleas Thoma, Felmont F. Eaves

Bibliographic record

VenueAesthetic Surgery Journal · 2017
Typeletter
Languageen
FieldEconomics, Econometrics and Finance
TopicHealth Systems, Economic Evaluations, Quality of Life
Canadian institutionsMcMaster UniversityImpact
Fundersnot available
KeywordsMedicineOutcome (game theory)Mathematical economics

Abstract

fetched live from OpenAlex

There are multiple factors that we should consider in any critical appraisal of a study. In this EBM Hub, we will use the study by Gama et al to discuss the choice of outcomes in aesthetic surgery.1 Whenever designing a study protocol − or even when reviewing the cases in your own practice − the choice of which outcome (or outcomes) to measure is critical. Outcome is an important component of the PICOT framework. This framework identifies the Population, Intervention, Comparative intervention, Outcomes and Time horizon of the study.2 We encourage all investigators to adopt the PICOT framework routinely. For any given intervention (ie, the treatment, or what you do), several types of outcomes (ie, what happens as a result) can be evaluated. Seems simple enough . . . right? You do something, something happens, and you measure that. Bada bing, bada boom. It’s Miller Time. OK, so it’s not really that simple. It turns out that there are lots of potential outcomes you can select for any particular intervention. Outcomes can be subjective (“how do you feel about your experience in the recovery room?”) or objective (“how long is the scar?”). Outcomes can have different perspectives from different vantage points, depending on who is measuring them or benefiting from them. Outcomes can be all or none (binary), such as “you have a hematoma or you don’t,” but outcomes can also be incremental, such as how many cubic centimeters (cc) of blood were lost. How you measure that outcome is as important as the outcome you are measuring. The how is the outcome metric. Outcome metrics can take all kinds of forms, including scales such as a Likert scale (1-10), an objective physical measurement (cc’s, blood pressure, nipple to inframammary fold distance), photographic analysis, or a survey, just to name a few. Outcomes can be viewed from a patient, provider, system, or economic perspective, but many outcomes may cross perspectives. Surgeon related outcomes might include things such as how often you successfully performed a step in the procedure, how much your neck hurts after that rhinoplasty case, or the time of the procedure, as studied by our authors here. Time, however, might also be seen as a system-related outcome (ie, effects preoperative flow demands, staffing, and operating room schedule). Time can also relate to an economic outcome, as (we can’t resist saying) time is money. However, the outcome research movement in the last 30 years has emphasized the importance of measuring outcomes that are important to patients rather than to physicians. One such article addressed the issue of outcomes in aesthetic surgery.3 Outcomes that are important to patients can be measured in a number of ways by different evaluators, but a big part of the outcomes research trends is in the use of patient-reported outcomes (PROs). PROs could include pain (either how much pain from surgery or how much pain relief after surgery), patient’s ratings of their aesthetic results, or other improvements related to quality of life (eg, improvement in their sex life or self-confidence). While some of the outcomes sought from patients are well defined, for example how many pain tablets were taken to be comfortable or how many patients experienced nausea, others seem “fuzzy,” such as self-confidence. However, in the last few years there has been a significant effort related to validating metrics for assessing these and many other types of patient reported outcomes. The BREAST-Q, BODY-Q, and FACE-Q are examples of patient reported outcome instruments that were methodically developed and validated and are totally applicable to aesthetic surgery.4 In the current study, the authors wanted to assess the outcome resulting from different suture methods in repairing diastasis recti.1 They prospectively randomized 30 patients into 3 groups. The control group underwent a two-layer repair. The two intervention groups had either a single layer of 2-0 nylon running suture (group 1) or a single layer of a barbed suture (group 2). The authors choose 2 outcomes to assess. The first was the time of repair of the diastasis and the second was diastasis recurrence, both of which were chosen to assess the “efficacy” of the repair. They found a 21-minute difference in closure between control and group 1 and a 22 minute difference between the control group and group 2. In order to assess recurrence, they performed a clinical assessment and postoperative ultrasound at 2 points postoperatively (3 weeks and 6 months), which they compared to preoperative ultrasound. Although the authors executed a good study in many ways, we feel that the first chosen outcome, time to repair the diastasis, has an overly narrow scope by representing primarily the perspective of the surgeon. We have to ask . . . how does this impact the patient? Does it affect their experience in the recovery room? Does it have an effect on their pain? Are they back at work quicker? Is it safer? Does it give a better result, worse result, or the same result? We feel that understanding how the different methods of plication effects patients would be an important outcome class to study in this population. If time is the outcome, as was chosen here, what effect did the difference in time have, other than on the surgeon? For example, it would be good to know if the intervention reduced cost to the system or to the patient or increased productivity in the operating room (eg, additional cases scheduled). In choosing outcomes, it is not only important to measure the right thing, but to determine how much of a difference is important. The interpretation of this difference is not routinely understood. We adopt novel procedures in which the difference from a prevailing procedure is found to be both statistically and clinically significant. (Who cares if a new procedure is found to be statistically significant in some outcome if we do not consider this difference to be clinically important?) This brings us to the concept of minimal clinically important difference (MCID). The MCID is the difference in outcomes that might make it worth providing the treatment, or for changing from one treatment to another. Let’s think through an example. Say treatment method A has a 30% incidence of seroma. If treatment B had a 5% incidence of seroma, or an 83% reduction, you’d probably say that’s an important difference. But what if it was a 5% or a 3% or a 1% difference? Would you say that such a small amount of difference is important enough to make a significant change in your technique (especially if treatment B is more expensive, painful, or takes longer)? In the author’s study, the difference in plication time between the control group and group 1 was 21 minutes, which represents a 60% reduction in plication time. On one hand, that seems impressive. However, if we look at the total operative time between the groups, the reduction was only 13.1% difference, which seems perhaps less important. From the patient perspective, however, we should ask if 21 minutes makes any difference at all. In terms of economic or system perspective, it might be that 21 minutes is a minimally important difference, although it would take an actual economic analysis to determine that. Before embarking on a study, it is important to not only determine the outcome you are measuring, but the MCID (in another Hub we’ll discuss how the MCID relates to a power analysis). The second outcome the authors chose, recurrence of diastasis, is an important outcome to study. A recurrence would significantly impact the patient, for instance by degrading their aesthetic result or causing them to undergo additional surgery. In studying the patients, the authors queried the patients and performed a physical exam to determine if there was a recurrent diastasis, but as the ultimate measurement they used an ultrasonic assessment to determine if the patient had a recurrence. An ultrasonic measurement rather than clinical exam or patient-reported measure of recurrence represents a surrogate metric. Surrogate metrics in a study might relate to the actual outcome we’d like to measure, but are generally easier (or cheaper or quicker) to measure. Let’s consider an example of surrogate measurements and how they are used (or abused). When trying to prevent cardiovascular disease, we are trying to prevent the major events that impact patients (for example, heart attack, sudden death, or stroke). However, given the relatively low rate of these events on an annual basis, it takes years (maybe decades), and many millions of dollars to study if a new intervention actually reduces these events. So, if you are trying to market (oops, we mean academically study) a new drug, the temptation is to not measure how many heart attacks were prevented or decreased sudden deaths (takes a long time and expensive − and your drug might fail), but to measure a reduction in blood pressure, lipid levels, or other laboratory tests (which are fast and cheaper and requires a fraction of the study subjects). We must ask ourselves if ultrasonic evaluation is the best way to measure diastasis recurrence, and if this is a surrogate metric. The actual outcome we’d like to know is whether a recurrence of the diastasis developed that caused the patient symptoms or deterioration in their results (ie, was it clinically significant). However, in the study patients, it appears that neither the authors nor the patients appreciated a clinical recurrence in any of the patients, rather the recurrence was diagnosed only by ultrasound. (This corroborates a previous study that failed to show a relationship between ultrasonic recurrence of diastasis recti and the development of symptoms.5) Thinking back to the concept of MCID, we also would ask, “If we use ultrasound for our diagnosis, how much of a difference is significant?” We would contend that ultrasonic separation of the muscles (a surrogate measurement of clinically appreciable separation) of only about a centimeter does not appear to be clinically significant. What would be interesting to know is how much of a difference correlates with a change in aesthetic results. It behooves investigators to be familiar with the most important or critical outcomes of the intervention they are investigating (Table 1). These outcomes should be placed in a hierarchical order. The top should be chosen as the primary outcome. The remaining ones should be the chosen as secondary outcomes. The sample size calculation is based on the primary outcome. Although more than one outcome could be considered primary, this would require a sample size which may be prohibitive as the power of the study must satisfy both outcomes. If the study had included patient quality of life as the primary outcome and the time to complete the repair of diastasis recti, clinical recurrence, and cost of each procedure, as secondary outcomes, the article would have had a greater impact. Choosing Outcomes: Our Recommendations MCID, minimal clinically important difference; PROs, patient-reported outcomes. Choosing Outcomes: Our Recommendations MCID, minimal clinically important difference; PROs, patient-reported outcomes. The authors have no conflict of interests to disclose related to the content of this article. The authors received no financial support for the research, authorship, and publication of this article.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.008
metaresearch head score (Gemma)0.072
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.053
Threshold uncertainty score0.046

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0080.072
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0040.005
Scholarly communication0.0060.005
Open science0.0010.003
Research integrity0.0530.049
Insufficient payload (model declined to judge)0.0100.006

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.436
GPT teacher head0.433
Teacher spread0.003 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2017
Admission routes1
Has abstractno

Explore more

Same venueAesthetic Surgery JournalSame topicHealth Systems, Economic Evaluations, Quality of LifeFrench-language works237,207