The Youth Fitness International Test (YFIT) battery for monitoring and surveillance among children and adolescents: A modified Delphi consensus project with 169 experts from 50 countries and territories
Bibliographic record
Abstract
BACKGROUND: Physical fitness in childhood and adolescence is associated with a variety of health outcomes and is a powerful marker of current and future health. However, inconsistencies in tests and protocols limit international monitoring and surveillance. The objective of the study was to seek international consensus on a proposed, evidence-informed, Youth Fitness International Test (YFIT) battery and protocols for health monitoring and surveillance in children and adolescents aged 6-18 years. METHODS: We conducted an international modified Delphi study to evaluate the level of agreement with a proposed, evidence-based, YFIT of core health-related fitness tests and protocols to be used worldwide in 6- to 18-year-olds. This proposal was based on previous European and North American projects that systematically reviewed the existing evidence to identify the most valid, reliable, health-related, safe, and feasible fitness tests to be used in children and adolescents aged 6-18 years. We designed a single-panel modified Delphi study and invited 216 experts from all around the world to answer this Delphi survey, of whom one-third are from low-to-middle income countries and one-third are women. Four experts were involved in the piloting of the survey and did not participate in the main Delphi study to avoid bias. We pre-defined an agreement of ≥80% among the expert participants to achieve consensus. RESULTS: We obtained a high response rate (78%) with a total of 169 fitness experts from 50 countries and territories, including 63 women and 61 experts from low- or middle-income countries/territories. Consensus (>85% agreement) was achieved for all proposed tests and protocols, supporting the YFIT battery, which includes weight and height (to compute body mass index as a proxy of body size/composition), the 20-m shuttle run (cardiorespiratory fitness), handgrip strength, and standing long jump (muscular fitness). CONCLUSION: This study contributes to standardizing fitness tests and protocols used for research, monitoring, and surveillance across the world, which will allow for future data pooling and the development of international and regional sex- and age-specific reference values, health-related cut-points, and a global picture of fitness among children and adolescents.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.124 | 0.077 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.002 |
| Science and technology studies | 0.003 | 0.003 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.009 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".