Performance Validity Testing in Patients with Substance Abuse in Addiction Care
Bibliographic record
Abstract
BACKGROUND: Low performances on neuropsychological tests are common in patients with substance use disorder (SUD), indicating potential cognitive impairments that may significantly impact treatment engagement and prognosis. While neuropsychological assessment is crucial for identifying these cognitive deficits, to date, data on the performance validity of individuals in addiction care is lacking. Performance validity testing (PVT) can be used to assess the accuracy of such test results. OBJECTIVES: This study examined the prevalence of suboptimal performance on different PVTs in a SUD inpatient population, their agreement in detecting poor performance validity, and their association with overall cognitive performance. METHODS: Retrospective data were analyzed from 172 SUD inpatients (2017-2024) in an addiction care clinic. Three PVTs were examined: the Visual Association Test-Extended (VAT-E), the Amsterdam Short-Term Memory test (ASTM), and the WAIS-IV Digit Span Age-Corrected Scaled Score (DS ACSS). Failure rates were calculated, and correlations between PVT outcomes and between the PVT measures and Montreal Cognitive Assessment (MoCA) were computed. RESULTS: Failure rates varied substantially across PVTs (from 1.3-36%). Agreement between PVTs was low (κ-values 0.019-0.397), with minimal correlations between ASTM, DS ACSS, and VAT-E scores. Weak to moderate positive correlations (ρ-values -0.024-0.403) were found between PVTs and the MoCA. CONCLUSIONS/IMPORTANCE: The variability in failure rates suggests that different PVTs may not measure the same construct. Possibly, the ASTM may be too challenging for many patients and DS ACSS failures may reflect lower intellectual abilities rather than true non-credible performance. This stresses the importance of selecting appropriate PVTs in addiction care settings to avoid misclassification and ensure valid neuropsychological assessments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".