Validity and Reliability of a Telehealth Physical Fitness and Functional Assessment Battery for Ambulatory Youth With and Without Mobility Disabilities: Observational Measurement Study
Bibliographic record
Abstract
BACKGROUND: Youth (age 15-24 years) with and without disability are not adequately represented enough in exercise research due to a lack of time and transportation. These barriers can be overcome by including accessible web-based assessments that eliminate the need for on-site visitations. There is no simple, low-cost, and psychometrically sound compilation of measures for physical fitness and function that can be applied to youth with and without mobility disabilities. OBJECTIVE: The first purpose was to determine the statistical level of agreement of 4 web-modified clinical assessments with how they are typically conducted in person at a laboratory (convergent validity). The second purpose was to determine the level of agreement between a novice and an expert rater (interrater reliability). The third purpose was to explore the feasibility of implementing the assessments via 2 metrics: safety and duration. METHODS: The study enrolled 19 ambulatory youth: 9 (47%) with cerebral palsy with various mobility disabilities from a children's hospital and 10 (53%) without disabilities from a university student population. Participants performed a battery of tests via videoconferencing and in person. The test condition (teleassessment and in person) order was randomized. The battery consisted of the hand grip strength test with a dynamometer, the five times sit-to-stand test (FTST), the timed up-and-go (TUG) test, and the 6-minute walk test (6MWT) either around a standard circular track (in person) or around a smaller home-modified track (teleassessment version, home-modified 6-minute walk test [HM6MWT]). Statistical analyses included descriptive data, intraclass correlation coefficients (ICCs), and Bland-Altman plots. RESULTS: The mean time to complete the in-person assessment was 16.9 (SD 4.8) minutes and the teleassessment was 21.1 (SD 5.9) minutes. No falls, injuries, or adverse events occurred. Excellent convergent validity was shown for telemeasured hand grip strength (right ICC=0.96, left ICC=0.98, P<.001) and the TUG test (ICC=0.92, P=.01). The FTST demonstrated good agreement (ICC=0.95, 95% CI 0.79-0.98; P=.01). The HM6MWT demonstrated poor absolute agreement with the 6MWT. However, further exploratory analysis revealed a strong positive correlation between the tests (r=0.83, P<.001). The interrater reliability was excellent for all tests (all ICCs>0.9, P<.05). CONCLUSIONS: This study suggests that videoconference assessments are convenient and useful measures of fitness and function among youth with and without disabilities. This paper presents operationalized teleassessment procedures that can be replicated by health professionals to produce valid and reliable measurements. This study is a first step toward developing teleassessments that can bypass the need for on-site data collection visitations for this age group. Further research is needed to identify psychometrically sound teleassessment procedures, particularly for measures of cardiorespiratory endurance or walking ability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.016 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".