Clinical utility of the Neurobehavioral Symptom Inventory validity scales to screen for symptom exaggeration following traumatic brain injury
Bibliographic record
Abstract
The purpose of this study was to examine the clinical utility of three recently developed validity scales (Validity-10, NIM5, and LOW6) designed to screen for symptom exaggeration using the Neurobehavioral Symptom Inventory (NSI). Participants were 272 U.S. military service members who sustained a mild, moderate, severe, or penetrating traumatic brain injury (TBI) and who were evaluated by the neuropsychology service at Walter Reed Army Medical Center within 199 weeks post injury. Participants were divided into two groups based on the Negative Impression Management scale of the Personality Assessment Inventory: (a) those who failed symptom validity testing (SVT-fail; n = 27) and (b) those who passed symptom validity testing (SVT-pass; n = 245). Participants in the SVT-fail group had significantly higher scores (p<.001) on the Validity-10, NIM5, LOW6, NSI total, and Personality Assessment Inventory (PAI) clinical scales (range: d = 0.76 to 2.34). Similarly high sensitivity, specificity, positive predictive power (PPP), and negative predictive (NPP) values were found when using all three validity scales to differentiate SVT-fail versus SVT-pass groups. However, the Validity-10 scale consistently had the highest overall values. The optimal cutoff score for the Validity-10 scale to identify possible symptom exaggeration was ≥19 (sensitivity = .59, specificity = .89, PPP = .74, NPP = .80). For the majority of people, these findings provide support for the use of the Validity-10 scale as a screening tool for possible symptom exaggeration. When scores on the Validity-10 exceed the cutoff score, it is recommended that (a) researchers and clinicians do not interpret responses on the NSI, and (b) clinicians follow up with a more detailed evaluation, using well-validated symptom validity measures (e.g., Minnesota Multiphasic Personality Inventory-2 Restructured Form, MMPI-2-RF, validity scales), to seek confirmatory evidence to support an hypothesis of symptom exaggeration.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".