Validity, Reliability, and Feasibility of Physical Literacy Assessments Designed for School Children: A Systematic Review
Bibliographic record
Abstract
BACKGROUND: While the burgeoning researcher and practitioner interest in physical literacy has stimulated new assessment approaches, the optimal tool for assessment among school-aged children remains unclear. OBJECTIVE: The purpose of this review was to: (i) identify assessment instruments designed to measure physical literacy in school-aged children; (ii) map instruments to a holistic construct of physical literacy (as specified by the Australian Physical Literacy Framework); (iii) document the validity and reliability for these instruments; and (iv) assess the feasibility of these instruments for use in school environments. DESIGN: This systematic review (registered with PROSPERO on 21 August, 2022) was conducted in accordance with the Preferred Reporting Items for Systematic Review and Meta-Analysis (PRISMA) statement. DATA SOURCES: Reviews of physical literacy assessments in the past 5 years (2017 +) were initially used to identify relevant assessments. Following that, a search (20 July, 2022) in six databases (CINAHL, ERIC, GlobalHealth, MEDLINE, PsycINFO, SPORTDiscus) was conducted for assessments that were missed/or published since publication of the reviews. Each step of screening involved evaluation from two authors, with any issues resolved through discussion with a third author. Nine instruments were identified from eight reviews. The database search identified 375 potential papers of which 67 full text papers were screened, resulting in 39 papers relevant to a physical literacy assessment. INCLUSION AND EXCLUSION CRITERIA: Instruments were classified against the Australian Physical Literacy Framework and needed to have assessed at least three of the Australian Physical Literacy Framework domains (i.e., psychological, social, cognitive, and/or physical). ANALYSES: Instruments were assessed for five aspects of validity (test content, response processes, internal structure, relations with other variables, and the consequences of testing). Feasibility in schools was documented according to time, space, equipment, training, and qualifications. RESULTS: Assessments with more validity/reliability evidence, according to age, were as follows: for children, the Physical Literacy in Children Questionnaire (PL-C Quest) and Passport for Life (PFL). For older children and adolescents, the Canadian Assessment for Physical Literacy (CAPL version 2). For adolescents, the Adolescent Physical Literacy Questionnaire (APLQ) and Portuguese Physical Literacy Assessment Questionnaire (PPLA-Q). Survey-based instruments were appraised to be the most feasible to administer in schools. CONCLUSIONS: This review identified optimal physical literacy assessments for children and adolescents based on current validity and reliability data. Instrument validity for specific populations was a clear gap, particularly for children with disability. While survey-based instruments were deemed the most feasible for use in schools, a comprehensive assessment may arguably require objective measures for elements in the physical domain. If a physical literacy assessment in schools is to be performed by teachers, this may require linking physical literacy to the curriculum and developing teachers' skills to develop and assess children's physical literacy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.072 | 0.293 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.011 | 0.014 |
| Bibliometrics | 0.015 | 0.013 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.005 | 0.005 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".