Review of criterion-referenced standards for cardiorespiratory fitness: what percentage of 1 142 026 international children and youth are apparently healthy?
Bibliographic record
Abstract
PURPOSE: To identify criterion-referenced standards for cardiorespiratory fitness (CRF); to estimate the percentage of children and youth that met each standard; and to discuss strategies to help improve the utility of criterion-referenced standards for population health research. METHODS: A search of four databases was undertaken to identify papers that reported criterion-referenced CRF standards for children and youth generated using the receiver operating characteristic curve technique. A pseudo-dataset representing the 20-m shuttle run test performance of 1 142 026 children and youth aged 9-17 years from 50 countries was generated using Monte Carlo simulation. Pseudo-data were used to estimate the international percentage of children and youth that met published criterion-referenced standards for CRF. RESULTS: Ten studies reported criterion-referenced standards for healthy CRF in children and youth. The mean percentage (±95% CI) of children and youth that met the standards varied substantially across age groups from 36%±13% to 95%±4% among girls, and from 51%±7% to 96%±16% among boys. There was an age gradient across all criterion-referenced standards where younger children were more likely to meet the standards compared with older children, regardless of sex. Within age groups, mean percentages were more precise (smaller CI) for younger girls and older boys. CONCLUSION: There are several CRF criterion-referenced standards for children and youth producing widely varying results. This study encourages using the interim international criterion-referenced standards of 35 and 42 mL/kg/min for girls and boys, respectively, to identify children and youth at risk of poor health-raising a clinical red flag.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.007 | 0.002 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".