MétaCan
Menu
Back to cohort
Record W4297857989 · doi:10.1101/2022.09.19.22279942

Concurrent and predictive validity of dynamic assessments of word reading in young children: A systematic review and meta-analysis

2022· review· en· W4297857989 on OpenAlexafffund
Emily Wood, Kereisha Biggs, Monika Molnar

Bibliographic record

VenuemedRxiv · 2022
Typereview
Languageen
FieldPsychology
TopicEducational and Psychological Assessments
Canadian institutionsToronto Rehabilitation InstituteUniversity of Toronto
FundersSocial Sciences and Humanities Research Council of CanadaNatural Sciences and Engineering Research Council of CanadaMinistry of Colleges and UniversitiesUniversity of Toronto
KeywordsConcurrent validityReading (process)Phonological awarenessPredictive validityPsychologyWord (group theory)LiteracyMeta-analysisCognitive psychologyDevelopmental psychologyLinguisticsPsychometricsMedicine

Abstract

fetched live from OpenAlex

Abstract Early evaluation of word reading skills is an important step in understanding and predicting children’s future literacy abilities. Traditionally, word reading evaluations are conducted using ‘static’ assessments (SA), which measure a child’s acquired knowledge and are prone to floor effects. Additionally, many of these tools are developed exclusively for English monolinguals, and therefore cannot be used equitably to evaluate the abilities of bilingual children. Dynamic assessment (DA), which evaluates the ability to learn a skill, is a potentially more equitable alternative. To establish that use of DAs is a valid alternative to traditional SAs, their concurrent agreement with gold standard SA measures and their predictive agreement with later word reading outcomes should be considered. In line with this, the primary objective of this systematic review and meta-analysis is to examine the concurrent and predictive validity of DAs of word reading skills. Two secondary objectives are (i) to address which types of word reading DAs (phonological awareness, sound-symbol knowledge, or decoding) demonstrate the strongest relationships with equivalent concurrent static measures and later word reading outcomes, and (ii) to consider for which populations, defined by language status (monolingual vs. bilingual vs. mixed) and reading status (typically developing vs. at-risk vs. mixed) these DAs are valid. Thirty-four studies from 32 papers were identified through searching 5 databases, and the grey literature. Included studies provided a correlation between a DA and concurrent SA, or a DA and a later word reading outcome measure. Regarding concurrent validity, we observed a strong relationship between DAs and SAs in general (r=.60); however, subgroup analyses indicate that DAs of decoding (r=.54) and phonological awareness (r=.73) measures demonstrate greater strength of correlation with their static counterparts, compared to DAs of sound-symbol knowledge (r=.34). In terms of predictive validity, we observed a similarly strong relationship between DAs and word reading outcome measures (r=.57), independently of the type of measure. Subgroup analyses conducted based on participant language status suggested that there are significant differences between mean effect sizes for monolingual, bilingual and mixed language groups in terms of DAs’ concurrent validity with SAs, but no significant differences for predictive validity with word reading outcome measures. There were also no significant differences between mean effect sizes for at-risk, typically developing, or mixed groups in terms of DAs concurrent validity with SAs or predictive validity with word reading outcome measures. Results provide preliminary evidence to suggest that DAs of phonological awareness and decoding skills are a valid alternative to SAs of equivalent constructs and are valid for the future prediction of word reading outcomes across population groups regardless of their language or reading status.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Insufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Meta-analysis · Consensus signal: none
GenreCandidate signal: Review · Consensus signal: Review
Teacher disagreement score0.765
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0020.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0070.001
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.196
GPT teacher head0.482
Teacher spread0.286 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designMeta-analysis
Domainnot available
GenreReview

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2022
Admission routes2
Has abstractyes

Explore more

Same venuemedRxivSame topicEducational and Psychological AssessmentsFrench-language works237,207