Psychometric properties and construct validity of PLAYself: a self-reported measure of physical literacy for children and youth
Bibliographic record
Abstract
PLAYself is a tool designed for self-description of physical literacy in children and youth. We examined the tool using both the Rasch model and Classical Test Theory to explore its psychometric properties. A random selection of 300 children aged 8–14 years (47.3% female) from a dataset of 8513 Canadian children were involved in the Rasch analysis. The 3 subscales of the measure demonstrated good fit to the Rasch model, satisfying requirements of unidimensionality, having good fit statistics (item and person fit residuals = –0.17–1.47) and internal reliability (Person Separation Index = 0.70–0.82), and a lack of item bias and problematic local dependency. In a separate comparable sample, 297 children also aged 8–14 years (53.9% female) completed the PLAYfun, Physical Self-Description Questionnaire (PSDQ), Physical Activities Measure-Revised (MPAM-R), a physical activity inventory (PLAYinventory), and repeated the PLAYself 7 days later. The tests with this sample confirmed test–retest reliability (intraclass correlation coefficient = 0.81–0.84), and convergent and construct validity consistent with contemporary physical literacy definitions. Overall, the PLAYself demonstrated robust psychometric properties, and is recommended for researchers and practitioners who are interested in assessing self-reported physical literacy. Novelty: The PLAYself is a self-reported measure of physical literacy This study validates the measure using the Rasch model and classical test theory The PLAYself was found to have strong psychometric properties
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".