Calibrating the Physical Activity Vital Sign to Estimate Habitual Moderate to Vigorous Physical Activity More Accurately in Active Young Adults: A Cautionary Tale
Bibliographic record
Abstract
The Physical Activity Vital Sign (PAVS) is a two-question assessment used to estimate habitual moderate to vigorous aerobic physical activity (MVPA). Previous studies have shown active adults cannot estimate the physical activity intensity properly. The initial purpose was to investigate the criterion validity of the PAVS for quantifying habitual MVPA in young adults meeting weekly MVPA guidelines (n = 140; 21 ± 3 years). A previously validated PiezoRx waist-worn accelerometer served as the criterion measure (wear time, 6.7 ± 0.6 days). All participants completed the PAVS once before wearing the PiezoRx. Standardized activity monitor validation procedures were followed. The PAVS (201 ± 142 min/week) underestimated (p < .001) MVPA compared to the PiezoRx (381 ± 155 min/week). To correct for this large error, the sample was divided into calibration model development (n = 70; 21 ± 3 years) and criterion validation (n = 70; 21 ± 3 years) groups. The PAVS score, age, gender, and body mass index outcomes from the development group were used to construct a multiple linear regression model-based calibrated PAVS (cPAVS) equation. In the validation group, the cPAVS was similar (p = .113; 352 ± 23 min/week) compared to accelerometry. Equivalence testing demonstrated the cPAVS, but not the PAVS, was equivalent to the PiezoRx. Despite achieving most statistical criteria, the PAVS and cPAVS still had high degrees of variability, preventing their use on an individual level. Alternative strategies are needed for the PAVS in an active young adult population. These results caution using the PAVS in active young adults and identify a case where obvious variabilities in accuracy conflict with statistically congruent results.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.089 | 0.222 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.002 | 0.005 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.004 | 0.002 |
| Research integrity | 0.005 | 0.008 |
| Insufficient payload (model declined to judge) | 0.001 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".