Calibrating the Physical Activity Vital Sign to Estimate Habitual Moderate to Vigorous Physical Activity More Accurately in Active Young Adults: A Cautionary Tale
Bibliographic record
Abstract
The Physical Activity Vital Sign (PAVS) is a two-question assessment used to estimate habitual moderate to vigorous aerobic physical activity (MVPA). Previous studies have shown active adults cannot estimate the physical activity intensity properly. The initial purpose was to investigate the criterion validity of the PAVS for quantifying habitual MVPA in young adults meeting weekly MVPA guidelines ( n = 140; 21 ± 3 years). A previously validated PiezoRx waist-worn accelerometer served as the criterion measure (wear time, 6.7 ± 0.6 days). All participants completed the PAVS once before wearing the PiezoRx. Standardized activity monitor validation procedures were followed. The PAVS (201 ± 142 min/week) underestimated ( p < .001) MVPA compared to the PiezoRx (381 ± 155 min/week). To correct for this large error, the sample was divided into calibration model development ( n = 70; 21 ± 3 years) and criterion validation ( n = 70; 21 ± 3 years) groups. The PAVS score, age, gender, and body mass index outcomes from the development group were used to construct a multiple linear regression model-based calibrated PAVS (cPAVS) equation. In the validation group, the cPAVS was similar ( p = .113; 352 ± 23 min/week) compared to accelerometry. Equivalence testing demonstrated the cPAVS, but not the PAVS, was equivalent to the PiezoRx. Despite achieving most statistical criteria, the PAVS and cPAVS still had high degrees of variability, preventing their use on an individual level. Alternative strategies are needed for the PAVS in an active young adult population. These results caution using the PAVS in active young adults and identify a case where obvious variabilities in accuracy conflict with statistically congruent results.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".