The risk–return trade‐off: Performance assessments and cognitive validation of inferences
Bibliographic record
Abstract
BACKGROUND AND AIMS: In educational measurement, performance assessments occupy a niche for offering a true-to-life format that affords the measurement of high-level cognitive competencies and the evidence to draw inferences about intellectual capital. However, true-to-life formats also introduce myriad complexities and can skew if not outright distort the accuracy of inferences. For validating claims about test-takers from performance assessments, the collection of evidence about response processes is a necessity of sufficient import that the validation process needs to be labelled a cognitive validation to ensure that the cognitive is not forgotten in the logic of the validation process. ANALYSIS AND EXAMPLE: Cognitive validation is described as a three-pronged process of (1) identifying the knowledge, skills, and attributes associated with the intellectual capital of interest, (2) selecting and/or developing tasks to elicit intellectual capital, and (3) collecting substantive empirical evidence of examinee response processes as part of the overall validity argument. This three-pronged process is illustrated using the American Institute of CPA's (2018) practice analysis, task-based simulations (TBSs), and use of think-aloud interviews to evaluate claims. CONCLUSIONS: Although cognitive laboratories and think alouds are used to measure distinct types of response processes as test-takers interact with performance assessments, both methods are among the best for obtaining direct but differential evidence from test-takers. The labour and cost of collecting this evidence are often not done or not done well by many testing programmes. However, for performance assessments to succeed in measuring what they purport to measure, the investment of cognitive validation must be made.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.380 | 0.815 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.007 | 0.005 |
| Science and technology studies | 0.002 | 0.024 |
| Scholarly communication | 0.014 | 0.029 |
| Open science | 0.004 | 0.011 |
| Research integrity | 0.008 | 0.008 |
| Insufficient payload (model declined to judge) | 0.007 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".