Reliability and validity of physical examination tests for the assessment of ankle instability
Bibliographic record
Abstract
INTRODUCTION: Clinicians rely on certain physical examination tests to diagnose and potentially grade ankle sprains and ankle instability. Diagnostic error and inaccurate prognosis may have important repercussions for clinical decision-making and patient outcomes. Therefore, it is important to recognize the diagnostic value of orthopaedic tests through understanding the reliability and validity of these tests. OBJECTIVE: To systematically review and report evidence on the reliability and validity of orthopaedic tests for the diagnosis of ankle sprains and instability. METHODS: PubMed, CINAHL, Scopus, and Cochrane databases were searched from inception to December 2021. In addition, the reference list of included studies, located systematic reviews, and orthopaedic textbooks were searched. All articles reporting reliability or validity of physical examination or orthopaedic tests to diagnose ankle instability or sprains were included. Methodological quality of the reliability and the validity studies was assessed with The Quality Appraisal for Reliability studies checklist and the Quality Assessment of Diagnostic Accuracy Studies-2 respectively. We identified the number of times the orthopaedic test was investigated and the validity and/or reliability of each test. RESULTS: Overall, sixteen studies were included. Three studies assessed reliability, eight assessed validity, and five evaluated both. Overall, fifteen tests were evaluated, none demonstrated robust reliability and validity scores. The anterolateral talar palpation test reported the highest diagnostic accuracy. Further, the anterior drawer test, the anterolateral talar palpation, the reverse anterior lateral drawer test, and palpation of the anterior talofibular ligament reported the highest sensitivity. The highest specificity was attributed to the anterior drawer test, the anterolateral drawer test, the reverse anterior lateral drawer test, tenderness on palpation of the proximal fibular, and the squeeze test. CONCLUSION: Overall, the diagnostic accuracy, reliability, and validity of physical examination tests for the assessment of ankle instability were limited. Physical examination tests should not be used in isolation, but rather in combination with the clinical history to diagnose an ankle sprain. Preliminary evidence suggests that the overall validity of physical examination for the ankle may be better if conducted five days after the injury rather than within 48 h of injury.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".