Reliability and validity of Canada???s Physical Activity Monitor for assessing trends
Bibliographic record
Abstract
PURPOSE: This investigation assessed the reliability and criterion validity of the Physical Activity Monitor, a telephone-interview adaptation of the Minnesota Leisure Time Physical Activity Questionnaire (MLTPAQ), which is currently used to assess trends in the Canadian population. METHODS: A sample of 512 people aged 18 yr and older was selected by random-digit dialing for telephone interviewing in the reliability study. The Monitor questions were administered twice, 3 wk apart. For the criterion validity study, a sample of 148 people aged 18-69 yr was selected at random from households. Participants completed the Monitor questions by telephone and an in-home step test to estimate maximum oxygen uptake. Another random sample of individuals aged 18-69 yr participated in a comparison study of the Monitor against the 1988 Campbell's Survey of Well-Being (CSWB) instrument. All studies were conducted in the vicinity of Toronto, Ontario. Spearman correlations controlling for age and sex were calculated as a measure of association for the reliability, validity, and comparison studies. Validity estimates were further adjusted for body mass index and physical activity demands of work and chores. RESULTS: The Monitor instrument produced reliable estimates of total energy expenditure (P=0.90, P<0.0001) with criterion validity of 0.36 (P<0.0001). The association between estimates of total energy expenditure derived from the Monitor and CSWB instruments was 0.77 (P<0.0001). CONCLUSION: The Physical Activity Monitor has acceptable test-retest reliability and criterion validity. The research also demonstrated that for the purpose of population monitoring a change in data collection mode-telephone interview versus self-administration in households-can yield reasonably comparable estimates from two adaptations of the MLTPAQ.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".