The Balance Evaluation Systems Test (BESTest) to Differentiate Balance Deficits
Bibliographic record
Abstract
BACKGROUND: Current clinical balance assessment tools do not aim to help therapists identify the underlying postural control systems responsible for poor functional balance. By identifying the disordered systems underlying balance control, therapists can direct specific types of intervention for different types of balance problems. OBJECTIVE: The goal of this study was to develop a clinical balance assessment tool that aims to target 6 different balance control systems so that specific rehabilitation approaches can be designed for different balance deficits. This article presents the theoretical framework, interrater reliability, and preliminary concurrent validity for this new instrument, the Balance Evaluation Systems Test (BESTest). DESIGN: The BESTest consists of 36 items, grouped into 6 systems: "Biomechanical Constraints," "Stability Limits/Verticality," "Anticipatory Postural Adjustments," "Postural Responses," "Sensory Orientation," and "Stability in Gait." METHODS: In 2 interrater trials, 22 subjects with and without balance disorders, ranging in age from 50 to 88 years, were rated concurrently on the BESTest by 19 therapists, students, and balance researchers. Concurrent validity was measured by correlation between the BESTest and balance confidence, as assessed with the Activities-specific Balance Confidence (ABC) Scale. RESULTS: Consistent with our theoretical framework, subjects with different diagnoses scored poorly on different sections of the BESTest. The intraclass correlation coefficient (ICC) for interrater reliability for the test as a whole was .91, with the 6 section ICCs ranging from .79 to .96. The Kendall coefficient of concordance among raters ranged from .46 to 1.00 for the 36 individual items. Concurrent validity of the correlation between the BESTest and the ABC Scale was r=.636, P<.01. LIMITATIONS: Further testing is needed to determine whether: (1) the sections of the BESTest actually detect independent balance deficits, (2) other systems important for balance control should be added, and (3) a shorter version of the test is possible by eliminating redundant or insensitive items. CONCLUSIONS: The BESTest is easy to learn to administer, with excellent reliability and very good validity. It is unique in allowing clinicians to determine the type of balance problems to direct specific treatments for their patients. By organizing clinical balance test items already in use, combined with new items not currently available, the BESTest is the most comprehensive clinical balance tool available and warrants further development.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".