MétaCan
Menu
Back to cohort
Record W1998804657 · doi:10.1037/spq0000001

The impact of probe variability on brief experimental analysis of reading skills.

2012· article· en· W1998804657 on OpenAlexaff
Sterett H. Mercer, Lauren Lestremau Harpole, Rachel R. Mitchell, Chandler McLemore, Christina Michelle Hardy

Bibliographic record

VenueSchool Psychology Quarterly · 2012
Typearticle
Languageen
FieldPsychology
TopicReading and Literacy Development
Canadian institutionsUniversity of British Columbia
FundersSociety for the Study of School Psychology
KeywordsFluencyPsychologyReplicateReading (process)Set (abstract data type)Intervention (counseling)Replication (statistics)Psychological interventionReliability (semiconductor)Developmental psychologyStatisticsMathematics educationComputer scienceMathematicsLinguistics

Abstract

fetched live from OpenAlex

The purpose of this study was to examine the impact of probe variability on the ability to replicate results in brief experimental analysis (BEA) of reading. In the first phase of the study, 41 first- and second- grade students completed 16 oral reading fluency probes. Calculations of probe difficulty were used to identify Low and High Variability probe sets. In the second phase of the study, the performance of 40 second- through fifth-grade students during two reading interventions was compared. The best-performing intervention for each student in the initial trial was replicated during a second trial for only 43% of students regardless of probe variability. The best-performing intervention was replicated for 60% of students when average performance across two trials was compared. Rules for determining the best-performing intervention in academic BEA should consider the standard error of measurement (SEM) for the probe set to be used, the reliability for absolute decisions using the probe set, and the number of replications relative to SEM needed to adequately demonstrate experimental control.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.043
metaresearch head score (Gemma)0.331
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.957
Threshold uncertainty score0.226

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0430.331
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0010.001
Scholarly communication0.0010.002
Open science0.0020.003
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.016
GPT teacher head0.383
Teacher spread0.367 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designObservational
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations5
Published2012
Admission routes1
Has abstractyes

Explore more

Same venueSchool Psychology QuarterlySame topicReading and Literacy DevelopmentFrench-language works237,207