Assessment of physical fitness during pregnancy: validity and reliability of fitness tests, and relationship with maternal and neonatal health – a systematic review
Bibliographic record
Abstract
Objectives: To systematically review studies evaluating one or more components of physical fitness (PF) in pregnant women, to answer two research questions: (1) What tests have been employed to assess PF in pregnant women? and (2) What is the validity and reliability of these tests and their relationship with maternal and neonatal health? Design: A systematic review. Data sources: PubMed and Web of Science. Eligibility criteria: Original English or Spanish full-text articles in a group of healthy pregnant women which at least one component of PF was assessed (field based or laboratory tests). Results: A total of 149 articles containing a sum of 191 fitness tests were included. Among the 191 fitness tests, 99 (ie, 52%) assessed cardiorespiratory fitness through 75 different protocols, 28 (15%) assessed muscular fitness through 16 different protocols, 14 (7%) assessed flexibility through 13 different protocols, 45 (24%) assessed balance through 40 different protocols, 2 assessed speed with the same protocol and 3 were multidimensional tests using one protocol. A total of 19 articles with 23 tests (13%) assessed either validity (n=4), reliability (n=6) or the relationship of PF with maternal and neonatal health (n=16). Conclusion: Physical fitness has been assessed through a wide variety of protocols, mostly lacking validity and reliability data, and no consensus exists on the most suitable fitness tests to be performed during pregnancy. PROSPERO registration number: CRD42018117554.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.007 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".