What Are the Levels of Evidence on Which We Base Decisions for Surgical Management of Lower Extremity Bone Tumors?
Bibliographic record
Abstract
BACKGROUND: Benign and malignant lower extremity primary bone tumors are among the least common conditions treated by orthopaedic surgeons. The literature supporting their surgical management has historically been in the form of observational studies rather than prospective controlled studies. Observational studies are prone to confounding bias, sampling bias, and recall bias. QUESTIONS/PURPOSES: (1) What are the overall levels of evidence of articles published on the surgical management of lower extremity bone tumors? (2) What is the overall quality of reporting of studies in this field based on the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) checklist? (3) What are the most common pitfalls in reporting that authors might improve on? METHODS: All studies describing the surgical management of lower extremity primary bone tumors from 2002 to 2012 were systematically reviewed. Two authors independently appraised levels of evidence. Quality of reporting was assessed with the STROBE checklist. Pitfalls in reporting were quantified by determining the 10 most underreported elements of research study design in the group of studies analyzed, again using the STROBE checklist as the reference standard. Of 1387 studies identified, 607 met eligibility criteria. RESULTS: There were no Level I studies, two Level II studies, 47 Level III studies, 308 Level IV studies, and 250 Level V studies. The mean percentage of STROBE points reported satisfactorily in each article as graded by the two reviewers was 53% (95% confidence interval, 42%-63%). The most common pitfalls in reporting were failures to justify sample size (2.2% reported), examine sensitivity (2.2%), account for missing data (9.8%), and discuss sources of bias (14%). Followup (66%), precision of outcomes (64%), eligibility criteria (55%), and methodological limitations (53%) were variably reported. CONCLUSIONS: Observational studies are the dominant evidence for the surgical management of primary lower extremity bone tumors. Numerous deficiencies in reporting limit their clinical use. Authors may use these results to inform future work and improve reporting in observational studies, and treating surgeons should be aware of these limitations when choosing among the various options with their patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.291 | 0.758 |
| Meta-epidemiology (narrow) | 0.002 | 0.003 |
| Meta-epidemiology (broad) | 0.014 | 0.017 |
| Bibliometrics | 0.026 | 0.018 |
| Science and technology studies | 0.004 | 0.008 |
| Scholarly communication | 0.021 | 0.021 |
| Open science | 0.013 | 0.008 |
| Research integrity | 0.015 | 0.009 |
| Insufficient payload (model declined to judge) | 0.009 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".