Base rates of non-credible performance in a post-secondary student sample seeking accessibility accommodations
Bibliographic record
Abstract
Objective: Performance Validity Tests (PVTs) have been used to identify non-credible performance in clinical, medicolegal, forensic, and, more recently, academic settings. The inclusion of PVTs when administering psychoeducational assessments is essential given that specific accommodation such as flexible deadlines and increased writing time can provide an external incentive for students without disabilities to feign symptoms. Method: The present study used archival data to establish base rates of non-credible performance in a sample of post-secondary students (n = 1045) who underwent a comprehensive psychoeducational evaluation for the purposes of obtaining academic accommodations. In accordance with current guidelines, non-credible performance was determined by failure on two or more freestanding or embedded PVTs. Results: 9.4% of participants failed at least two of the PVTs they were administered, of which 8.5% failed two PVTs, and approximately 1% failed three PVTs. Base rates of failure for specific PVTs ranged from 25% (b Test) to 11.2% (TOVA). Conclusions: The present study found a lower base rate of non-credible performance than previously observed in comparable populations. This likely reflects the utilization of conservative criteria in detecting non-credible performance to avoid false positives. By contrast, inconsistent base rates previously found in the literature may reflect inconsistent methodologies. These results further emphasize the importance of administering multiple PVTs during psychoeducational assessments. The implications of these findings can further inform clinicians administering assessments in academic settings and aid in the appropriate utilization of PVTs in psychoeducational evaluation to determine accessibility accommodations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".