Response
Bibliographic record
Abstract
Dear Editor-in-Chief, We thank Dr. Robinson and her colleagues for their interest in our work (1) and giving us the opportunity to respond to their letter. The pattern of results with fundamental movement skills was not different using standardized scores, or percentiles, as demonstrated in Table 2. We used percentiles as described in the Test of Gross Motor Development-2 analysis manual, not only because they allow comparisons between groups in change but also change relative to an age-matched population. With regard to adjustment for baseline values, these were not made because there were no significant baseline differences in mean gross motor quotient percentiles (p. 931) and, we used mixed models that adequately account for baseline differences. As for not adjusting for multiple analyses (e.g., Bonferroni), this decision was informed by Rothman (3) a leading epidemiologist and methodologist. “The theoretical basis for advocating a routine adjustment for multiple comparisons is the ‘universal null hypothesis’ that ‘chance’ serves as the first-order explanation for observed phenomena. This hypothesis undermines the basic premises of empirical research, which holds that nature follows regular laws that may be studied through observations. A policy of not making adjustments for multiple comparisons is preferable because it will lead to fewer errors of interpretation when the data under evaluation are not random numbers but actual observations on nature.” Similar to the pattern of results for standardized scores and percentile scores on the global fundamental movement skills indicators (i.e., GMQ, locomotor skills, object control skills), there were no differences in pattern of results when relative group differences were compared with raw scores versus percentile scores, on the individual locomotor skills and object control skills. Wanting to identify what set of skills might be driving the differences between groups, we chose to examine and present raw scores at both baseline and 6 months in Figures 3A and 3B. We hoped to reflect absolute performance for ease of comparison across studies, while still illustrating group differences at each time point. Visually comparing the unprocessed raw scores (Figs. 3A and 3B), with age-matched and sex-matched standardized scores (determined using tables of normative data (non-Canadian) reported in Table 2, is not appropriate. We did not intend to imply this and apologize for any confusion our presentation of data might have caused. Albeit normative data are not available for Canadian children as indicated, the TGMD-2 scoring guide indicates the GMQ to be the “best measure of an individual’s gross motor ability” and since both intervention and control groups are being compared to these percentiles, we believe our relative group differences are still meaningful. We agree with the need for more consistency in the TGMD score reporting and are hopeful that the updated tool and scoring rubric is more specific and directive. We agree with the suggestion to follow the CONSORT statement and we believe our study was conducted in line with the standards in the statement, as reflected in the primary article from this study (2), referenced in our methods. Given word limitations for MSSE and a request to condense the primary article, we were forced to include only the most pertinent information. Kristi B. Adamo School of Human Kinetics Healthy Active Living and Obesity Research Group Children’s Hospital of Eastern Ontario Research Institute Pediatrics, Faculty of Medicine University of Ottawa Ottawa, CANADA Shanna Wilson Healthy Active Living and Obesity Research Group Children’s Hospital of Eastern Ontario Research Institute Ottawa, CANADA Alysha L. J. Harvey Kimberly P. Grattan School of Human Kinetics University of Ottawa Ottawa, CANADA Patti-Jean Naylor School of Exercise Sciences, Physical & Health Education University of Victoria Victoria, CANADA Viviene A. Temple School of Exercise Sciences, Physical & Health University of Victoria Victoria, CANADA Gary S. Goldfield Healthy Active Living and Obesity Research Group Children’s Hospital of Eastern Ontario Research Institute Pediatrics, Faculty of Medicine University of Ottawa Ottawa, CANADA
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.038 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.014 | 0.015 |
| Insufficient payload (model declined to judge) | 0.112 | 0.079 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".