The Autonomy-validity Dilemma in Mechanical Prediction Procedures: The Quest for a Compromise
Bibliographic record
Abstract
A robust finding in psychological research is that combining information with a mechanical rule results in more valid predictions than combining information holistically in the mind. Nevertheless, information is typically combined holistically in practice, resulting in suboptimal predictions and decisions. Earlier research showed that decision makers are more likely to use mechanical prediction procedures when they retain autonomy in the decision-making process. However, it remains largely unknown how different autonomy-enhancing features affect predictive validity. Therefore, in two pre-registered studies (total N = 342), we investigated if and how prediction procedures can be designed such that they satisfy decision-makers’ autonomy needs and acceptance without reducing predictive validity. Based on archival application data from a university admission procedure, participants predicted applicants’ first-year GPA and chance of dropout. The results of Bayesian analyses showed that participants preferred prediction procedures in which they retained autonomy by choosing consistent predictor weights of a mechanical rule or by holistically adjusting the predictions of an optimal regression model. In general, these prediction procedures resulted in slightly higher predictive validity compared to fully holistic prediction. Providing participants with predictor validity information slightly increased predictive validity when participants could choose predictor weights, but not when making holistic predictions or adjusting optimal model predictions. Giving decision makers a role in designing mechanical rules through choosing weights based on explicit predictive validity information could help promote the implementation and validity of mechanical prediction in practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.127 | 0.351 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.002 | 0.019 |
| Scholarly communication | 0.008 | 0.018 |
| Open science | 0.004 | 0.007 |
| Research integrity | 0.005 | 0.007 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".