Concurrent and predictive validity of dynamic assessments of word reading in young children: A systematic review and meta-analysis
Bibliographic record
Abstract
Abstract Early evaluation of word reading skills is an important step in understanding and predicting children’s future literacy abilities. Traditionally, word reading evaluations are conducted using ‘static’ assessments (SA), which measure a child’s acquired knowledge and are prone to floor effects. Additionally, many of these tools are developed exclusively for English monolinguals, and therefore cannot be used equitably to evaluate the abilities of bilingual children. Dynamic assessment (DA), which evaluates the ability to learn a skill, is a potentially more equitable alternative. To establish that use of DAs is a valid alternative to traditional SAs, their concurrent agreement with gold standard SA measures and their predictive agreement with later word reading outcomes should be considered. In line with this, the primary objective of this systematic review and meta-analysis is to examine the concurrent and predictive validity of DAs of word reading skills. Two secondary objectives are (i) to address which types of word reading DAs (phonological awareness, sound-symbol knowledge, or decoding) demonstrate the strongest relationships with equivalent concurrent static measures and later word reading outcomes, and (ii) to consider for which populations, defined by language status (monolingual vs. bilingual vs. mixed) and reading status (typically developing vs. at-risk vs. mixed) these DAs are valid. Thirty-four studies from 32 papers were identified through searching 5 databases, and the grey literature. Included studies provided a correlation between a DA and concurrent SA, or a DA and a later word reading outcome measure. Regarding concurrent validity, we observed a strong relationship between DAs and SAs in general (r=.60); however, subgroup analyses indicate that DAs of decoding (r=.54) and phonological awareness (r=.73) measures demonstrate greater strength of correlation with their static counterparts, compared to DAs of sound-symbol knowledge (r=.34). In terms of predictive validity, we observed a similarly strong relationship between DAs and word reading outcome measures (r=.57), independently of the type of measure. Subgroup analyses conducted based on participant language status suggested that there are significant differences between mean effect sizes for monolingual, bilingual and mixed language groups in terms of DAs’ concurrent validity with SAs, but no significant differences for predictive validity with word reading outcome measures. There were also no significant differences between mean effect sizes for at-risk, typically developing, or mixed groups in terms of DAs concurrent validity with SAs or predictive validity with word reading outcome measures. Results provide preliminary evidence to suggest that DAs of phonological awareness and decoding skills are a valid alternative to SAs of equivalent constructs and are valid for the future prediction of word reading outcomes across population groups regardless of their language or reading status.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.007 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".