Examining the validity of the use of ratio IQs in psychological assessments
Bibliographic record
Abstract
Intelligence tests are amongst the most used psychological assessments, both in research and clinical settings. To avoid missing data points, for participants who cannot complete Intelligence tests normed for their age, ratio IQ scores (RIQ) are routinely computed and used as a proxy of IQ. Here, we use the case of autism to examine the validity of this widely used, yet never scientifically validated, practice. We examine the differences between standard full-scale IQ (FSIQ) and RIQ. Data was extracted from four databases in which age, FSIQ scores and subtests raw scores (from which RIQ scores could be calculated) were available for 16,751 autistic participants between 2 and 18 years old. The Intelligence tests included were the MSEL (N = 12,033), DAS-II early years (N = 1270), DAS-II school age (N = 2848), WISC-IV (N = 471) and WISC-V (N = 129). RIQs were computed for each participant as well as the discrepancy (DSC) between RIQ and FSIQ. We performed a multiple linear regression model to assess the effects of age and FSIQ on DSC for each IQ test. Participants at the extremes of the FSIQ distribution tended to have a greater DSC than participants with average FSIQ. Furthermore, age significantly predicted the DSC, with RIQ superior to FSIQ for younger participants while the opposite was found for older participants. Similar results were found in secondary analyses including typically developing children. These results question the validity of the RIQ as an alternative scoring method, especially for individuals at the extremes of the normal distribution, for whom RIQs are most often employed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".