Validity of the international physical activity questionnaire short form (IPAQ-SF) in adults : a systematic review
Bibliographic record
Abstract
Background \nAccording to the World Health Organization, sedentary lifestyle is a worldwide phenomenon nowadays and it is associated with many chronic diseases. Thus the promotion of active lifestyle is a public health priority. “The International Physical Activity Questionnaire Short Form” (IPAQ-SF) is recommended as an efficient and cost-effective tool for physical activity (PA) assessment. The validity of IPAQ-SF has been studied with inconsistent results within different age groups. This systematic review is aimed at analyzing and synthesizing validation tests of IPAQ-SF in adults. \n \nMethods \nPublished articles were searched with keywords of “International Physical Activity Questionnaire Short Form”, “IPAQ-SF”, “validation”, “validity”, and “adults” in bibliographic database “PubMed”, “Google scholar”, “Chinese National Knowledge Infrastructure” and “Wanfang Database”. Inclusion and exclusion criteria were applied and the results were extracted and summarized. \n \nResults \nTwelve studies published from January 2006 to January 2015 were included in this systematic review. Four studies were from North America (USA 2, Canada 1, Mexico 1), 1 study from South America (Brazil), 3 studies from Asia (Japan, Hong Kong, China), 2 studies from Europe (Norway and Greece), and 2 studies from Africa (Nigeria). Sample size ranged from 102 to 629 and subjects were aged from 23 to 51. The 12 papers included 14 validation analyses, using gold standards of objective movement assessments (6), objective fitness assessments (6), double labeled water (1) and physical activity log (1). \n \nFor the total PA score, the range of Spearman correlation r were -0.02 to 0.38 (mean 0.21) overall, 0.12 to 0.38 (mean 0.24) for studies using movement-assessment tools as the gold standard, and -0.02 to 0.36 (mean 0.19) for studies using fitness-assessing tools as the gold standard. None of the studies achieved the acceptable level of Spearman correlation r of 0.5. \n \nFor the specific PA score, higher Spearman correlation r was obtained for vigorous PA score (range -0.26 to 0.47, mean 0.24) than moderate PA score (range -0.09 to 0.5, mean 0.13), and walking PA score (range -0.05 to 0.17, mean 0.09). \n \nFor the accuracy of IPAQ-SF, only 5 papers reported the absolute difference of metabolic equivalent task (MET) between IPAQ-SF and gold standard. IPAQ-SF over-estimated MET in 4 studies, with “IPAQ- SF MET/ Criterion MET” ratio ranged from 1.11 to 6.41; IPAQ-SF under-estimated MET in 1 study, with an “IPAQ-SF MET/ Criterion MET” ratio of 0.52. \n \nConclusions \nThe validity of IPAQ-SF as measured by Spearman correlation r of 12 studies was low to moderate at best. Higher validity was obtained for vigorous PA than moderate PA and walking. IPAQ-SF also tended to overestimate MET. Given the difficulty of measuring PA objectively, IPAQ-SF is still useful in estimating physical activity, especially in large-scale studies. Future research should investigate the validity of IPAQ-SF in various subgroups (e.g. by age, sex, education level) to identify those with higher validity. Improved versions of IPAQ-SF and new scales could also be developed and examined.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.074 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.008 | 0.008 |
| Bibliometrics | 0.010 | 0.012 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".