Abstract: Meta-Analyses in Plastic Surgery: Can We Trust Their Results?
Bibliographic record
Abstract
PURPOSE: The objectives of this manuscript is to assess the overall quality of meta-analyses in plastic surgery from 2007–2017, assess whether there has been an improvement in quality over time, and evaluate variables that may be associated with scientific quality. METHODS: A systematic review of meta-analyses was undertaken using a computerized search of Medline, Embase, Cochrane Database for Systematic Reviews. Articles from seven plastic surgery journals published between the years 2007 to 2017 were included. Publication descriptors (author, year, country of publication), methodological and statistical methods were extracted. Each article was then assessed using the A Measurement Tool to Assess Systematic Reviews (AMSTAR) instrument. RESULTS: A total of 67 studies were included. The number of meta-analyses increased consistently between 2007 and 2017 with the majority of studies coming from the United States. Most studies were outcome based, assessing a single intervention, from the journal Plastic & Reconstructive Surgery, pooled a mean of 21 primary studies (range: 2–134), and utilized a mean of 2465 patients (range: 44-14884). Most meta-analyses analyzed primary studies in the middle tiers of evidence levels (II to IV), with a small percentage analyzing randomized controlled trials (16.4%). Random effect modeling was most commonly used (47.8%) and meta-analyses generally had positive (82.1%) and significant results (74.6%). Meta-analyses evaluated clinical (80.6%), methodological (65.6%), and statistical heterogeneity (50.7%) variably in terms of appropriateness and a substantial portion did not acknowledge or report methodological (7.5%) and statistical heterogeneity (25.4%). AMSTAR scores ranged between two and ten, with a mean of 6.7 out of 11. AMSTAR scores were correlated with year of publication (p=0.04, R=0.25). Multivariable linear analysis indicated that more recent studies, studies that included a rationale for statistical pooling, and studies that properly managed methodological heterogeneity were correlated with higher AMSTAR scores (r=0.66, p<0.01). CONCLUSION: The quality and number of meta-analyses have increased; however, despite an improvement in quality, the overall quality of most meta-analyses remains low. Meta-analyses should utilize proper data pooling methods and account for clinical heterogeneity appropriately. Readers, authors, reviewers, and journal editors should utilize validated instruments to evaluate meta-analysis to uphold methodological integrity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.078 | 0.264 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.012 | 0.005 |
| Bibliometrics | 0.001 | 0.006 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.004 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.025 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".