Has the Level of Evidence of Podium Presentations at the Musculoskeletal Tumor Society Annual Meeting Changed Over Time?
Bibliographic record
Abstract
BACKGROUND: Level of evidence (LOE) framework is a tool with which to categorize clinical studies based on their likelihood to be influenced by bias. Improvements in LOE have been demonstrated throughout orthopaedics, prompting our evaluation of orthopaedic oncology research LOE to determine if it has changed in kind. QUESTIONS/PURPOSES: (1) Has the LOE presented at the Musculoskeletal Tumor Society (MSTS) annual meeting improved over time? (2) Over the past decade, how do the MSTS and Orthopaedic Trauma Association (OTA) annual meetings compare regarding LOE overall and for the subset of therapeutic studies? METHODS: We reviewed abstracts from MSTS and OTA annual meeting podium presentations from 2005 to 2014. Three independent reviewers evaluated a total of 1222 abstracts for study type and LOE; there were 577 abstracts from MSTS and 645 from OTA. Changes in the distributions of study type and LOE over time were evaluated by Pearson chi-square test. RESULTS: There was no change over time in MSTS LOE for all study types (p = 0.13) and therapeutic (p = 0.36) study types during the reviewed decade. In contrast, OTA LOE increased over this time for all study types (p < 0.01). The proportion of Level I therapeutic studies was higher at the OTA than the MSTS (3% [14 of 413] versus 0.5% [two of 387], respectively), whereas the proportion of Level IV studies was lower at the OTA than the MSTS (32% [134 of 413] versus 75% [292 of 387], respectively) during the reviewed decade. The proportion of controlled therapeutic studies (LOE I through III) versus uncontrolled studies (LOE IV) increased over time at OTA (p < 0.021), but not at MSTS (p = 0.10). CONCLUSIONS: Uncontrolled case series continue to dominate the MSTS scientific program, limiting progress in evidence-based clinical care. Techniques used by the OTA to improve LOE may be emulated by the MSTS. These techniques focus on broad participation in multicenter collaborations that are designed in a comprehensive manner and answer a pragmatic clinical question.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.304 | 0.690 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.006 | 0.009 |
| Bibliometrics | 0.026 | 0.024 |
| Science and technology studies | 0.002 | 0.005 |
| Scholarly communication | 0.013 | 0.011 |
| Open science | 0.005 | 0.005 |
| Research integrity | 0.006 | 0.004 |
| Insufficient payload (model declined to judge) | 0.007 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".