Level of Clinical Evidence Presented at the Arthroscopy Association of North America Annual Meeting Over 10 Years (2006‐2015)
Bibliographic record
Abstract
PURPOSE: To evaluate any trends in the level of clinical evidence in the papers presented at the Arthroscopy Association of North America (AANA) annual scientific meetings from 2006 to 2015. METHODS: The online abstracts of the paper presentations presented at the AANA meetings were independently evaluated by 2 reviewers (664 total presentations). The reviewers independently screened these results for clinical studies and graded their level of evidence from Level I (i.e., randomized trials) to IV (i.e., case series) based on the American Academy of Orthopaedic Surgeons classification system. RESULTS: Five hundred thirteen presentations met the inclusion criteria and were evaluated. Overall, 16% of the presentations were Level I evidence, 15% were Level II, 26% were Level III, and 43% were Level IV. We observed a significant non-random improvement in the level of evidence of presentations at the AANA meetings (P ≤ .001) between 2006 and 2015. In particular, the percentage of papers with Level IV evidence presented significantly decreased (P ≤ .001) and the percentage of papers with Level III evidence increased (P = .004) over the study period. CONCLUSIONS: Statistical trends show that the influence of evidence-based medicine in orthopaedics has had a positive impact on the quality of research presented at the AANA meetings. LEVEL OF EVIDENCE: Level IV, review of abstracts of Level I to Level IV evidence.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.025 | 0.182 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.005 | 0.005 |
| Bibliometrics | 0.012 | 0.012 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.010 | 0.005 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.008 | 0.003 |
| Insufficient payload (model declined to judge) | 0.047 | 0.012 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".