Flawed Studies of SIMS’s Diagnostic Accuracy by Teams of Puente-López and Capilla Ramírez
Bibliographic record
Abstract
Background: The teams of Puente-López and Capilla Ramírez evaluated diagnostic accuracy of the Structured Inventory of Malingered Symptomatology (SIMS), a test often used to assess malingering by persons injured in motor vehicle accidents (MVAs). Yet all SIMS items represent legitimate medical symptoms, and more than 50% of them are those experienced by severely injured motorists, but they are fallaciously scored as indicative of malingering. Thus, more injured patients with more symptoms obtain higher SIMS scores for malingering. Method: The studies by Puente-López and by Capilla Ramírez were carried out on SIMS scores of injured motorists. The present article assesses the severity of their injuries, as documented by Puente-López and by Capilla Ramírez. Results and Discussion: The study by Capilla Ramírez’s team excluded patients with pathological results on physical examinations, or on X-Rays, EMG, and MRI: thus, only mildly injured motorists were included. The patients of Puente-López had signs of only a mild cervical whiplash. Almost none reported lower back pain or dizziness. Thus, both studies included patients with only mild symptoms that resulted in very low SIMS scores: they scored within the non-malingering range as defined by the SIMS manual. Their scores were below SIMS scores of healthy persons instructed to feign whiplash symptoms from an MVA. The teams of Capilla Ramírez and of Puente-López erroneously interpreted these results as demonstrating diagnostic accuracy of the SIMS for detection of malingering in injured motorists. Conclusions: The two studies of very mildly injured motorists fail to demonstrate “diagnostic accuracy of the SIMS” because the SIMS is mostly used by insurance contracted psychologists on more severely injured MVA patients (those with whiplash and post-concussion syndrome), i.e., those with more symptoms and thus, with higher SIMS scores that fallaciously classify them as “malingerers.”
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.012 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".