Cross-Validating the Atypical Response Scale of the TSI-2 in a Sample of Motor Vehicle Collision Survivors
Bibliographic record
Abstract
Abstract This study was designed to evaluate the utility of the Atypical Responses (ATR) scale of the Trauma Symptom Inventory – Second Edition (TSI-2) as a symptom validity test (SVT) in a medicolegal sample. Archival data were collected from a consecutive case sequence of 99 patients referred for neuropsychological evaluation following a motor vehicle collision. The ATR’s classification accuracy was computed against criterion measures consisting of composite indices based on SVTs and performance validity tests (PVTs). An ATR cutoff of ≥ 9 emerged as the optimal cutoff, producing a good combination of sensitivity (.35-.53) and specificity (.92-.95) to the criterion SVT, correctly classifying 71–79% of the sample. Predictably, classification accuracy was lower against PVTs as criterion measures (.26-.37 sensitivity at .90-.93 specificity, correctly classifying 66–69% of the sample). The originally proposed ATR cutoff (≥ 15) was prohibitively conservative, resulting in a 90–95% false negative rate. In contrast, although the more liberal alternative (≥ 8) fell short of the specificity standard (.89), it was associated with notably higher sensitivity (.43-.68) and the highest overall classification accuracy (71–82% of the sample). Non-credible symptom report was a stronger confound on the posttraumatic stress scale of the TSI-2 than that of the Personality Assessment Inventory. The ATR demonstrated its clinical utility in identifying non-credible symptom report (and to a lesser extent, invalid performance) in a medicolegal setting, with ≥ 9 emerging as the optimal cutoff. The ATR demonstrated its potential to serve as a quick (potentially stand-alone) screener for the overall credibility of neuropsychological deficits. More research is needed in patients with different clinical characteristics assessed in different settings to establish the generalizability of the findings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".