Bibliographic record
Abstract
Dear Editor-in-Chief, The concept of physical literacy (PL) and developing physically literate youth has become a topic of discussion among many researchers, across a plethora of disciplines. PL is often used to describe the characteristics of an individual who has the knowledge, physical competence, motivation, and confidence to be physically active for life (1). Measuring this construct will ultimately prove to be difficult, as the multifaceted nature of physically literacy is hard to capture in one assessment tool. Given our research group’s keen interest in PL and the measurement of PL, we were extremely interested to read about the Dragon Challenge, a dynamic assessment tool of physical competence for children 10 to 14 yr of age (2). Although the Dragon Challenge is not a comprehensive assessment of all domains of PL, the authors certainly provide a novel and interesting approach to measuring one of the core domains of the concept—physical or motoric competence. Unlike most movement skill assessments, the researchers state that the Dragon Challenge provides a comprehensive measurement of physical competence as multiple series of motor skills are assessed (i.e., simple, complex and combined) in an authentic environment, thereby ensuring the most accurate measure of PL. This is in opposition, the authors argue, to assessments that measure discrete skills in isolation, performed in static, limited environmental conditions. It is here that they identify a number of different assessments such as the Test of Gross Motor Development-2 (3), Bruninks–Oseretsky Test of Motor Proficiency (4), Movement Assessment Battery of Children-2 (5), and the Physical Literacy Assessment for Youth (PLAYfun) (6). Indeed, this represents a broad category of measures, some of which were developed for clinical purposes, whereas others were more consistent with assessment of motor competence in general populations. We contend that the authors have not accurately represented several of these measures and as a result overstate the novelity and uniqueness of their new measure. For instance, the PLAYfun tool is not simply a measure of movement competence. The tool also quantifies other domains of PL such as competence and confidence. In addition, although the authors imply that PLAYfun is a test that involves discrete skills, this is an oversimplification. The PLAYfun tool examines multiple facets of each of the 18 movement skills included in the battery. For example, when assessing running in a square, assessors look for not only proper running form but also the ability to pivot and speed of the movement (6). Even with a test such as the BOTMP, which focuses mainly on motor proficiency, skills are combined together. For instance, upper body limb coordination, wherein skills such as the dribble are tested, requires bilateral coordination of the hands (4). Although the Dragon Challenge certainly adds an important, alternate approach to the assessment of PL, we must be careful not to set up false propositions about the novelty or uniqueness of the approach. This tends to exaggerate differences between measures and may lead to unnecessary confusion among researcher and practitioners who are looking for measures to use. Laura St. John John Cairney INfant and Child Health (INCH) Laboratory Department of Family Medicine McMaster University Hamilton Ontario, CANADA
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.046 | 0.273 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.005 | 0.003 |
| Science and technology studies | 0.005 | 0.016 |
| Scholarly communication | 0.012 | 0.013 |
| Open science | 0.008 | 0.007 |
| Research integrity | 0.032 | 0.050 |
| Insufficient payload (model declined to judge) | 0.013 | 0.014 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".