Test-Retest Reliability of Home-Based Fitness Assessments Using a Mobile App (R Plus Health) in Healthy Adults: Prospective Quantitative Study
Bibliographic record
Abstract
BACKGROUND: Poor physical fitness has a negative impact on overall health status. An increasing number of health-related mobile apps have emerged to reduce the burden of medical care and the inconvenience of long-distance travel. However, few studies have been conducted on home-based fitness tests using apps. Insufficient monitoring of physiological signals during fitness assessments have been noted. Therefore, we developed R Plus Health, a digital health app that incorporates all the components of a fitness assessment with concomitant physiological signal monitoring. OBJECTIVE: The aim of this study is to investigate the test-retest reliability of home-based fitness assessments using the R Plus Health app in healthy adults. METHODS: A total of 31 healthy young adults self-executed 2 fitness assessments using the R Plus Health app, with a 2- to 3-day interval between assessments. The fitness assessments included cardiorespiratory endurance, strength, flexibility, mobility, and balance tests. The intraclass correlation coefficient was computed as a measure of the relative reliability of the fitness assessments and determined their consistency. The SE of measurement, smallest real difference at a 90% CI, and Bland-Altman analyses were used to assess agreement, sensitivity to real change, and systematic bias detection, respectively. RESULTS: The relative reliability of the fitness assessments using R Plus Health was moderate to good (intraclass correlation coefficient 0.8-0.99 for raw scores, 0.69-0.99 for converted scores). The SE of measurement and smallest real difference at a 90% CI were 1.44-6.91 and 3.36-16.11, respectively, in all fitness assessments. The 95% CI of the mean difference indicated no significant systematic error between the assessments for the strength and balance tests. The Bland-Altman analyses revealed no significant systematic bias between the assessments for all tests, with a few outliers. The Bland-Altman plots illustrated narrow limits of agreement for upper extremity strength, abdominal strength, and right leg stance tests, indicating good agreement between the 2 assessments. CONCLUSIONS: Home-based fitness assessments using the R Plus Health app were reliable and feasible in young, healthy adults. The results of the fitness assessments can offer a comprehensive understanding of general health status and help prescribe safe and suitable exercise training regimens. In future work, the app will be tested in different populations (eg, patients with chronic diseases or users with poor fitness), and the results will be compared with clinical test results. TRIAL REGISTRATION: Chinese Clinical Trial Registry ChiCTR2000030905; http://www.chictr.org.cn/showproj.aspx?proj=50229.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.016 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".