A classification-image-like method reveals strategies in 2afc tasks
Bibliographic record
Abstract
Despite decades of research, there is still uncertainty about how observers make even the simplest visual judgements, such as 2AFC decisions. Here we demonstrate a new method of using classification images to calculate "proxy decision variables" that estimate an observer's decision variables on individual trials. This provides a new way of investigating decision strategies. In Experiment 1, nine observers viewed two disks in Gaussian noise, to the left and right of fixation, and judged which had a contrast increment. The contrast increment was set to each observer's 70% threshold. On each trial we calculated the cross-correlation of the observer's classification image with the two disks, providing proxy decision variables. Using 10,000 such trials per observer we mapped the observer's decision space: we plotted the probability of the observer choosing the right-hand disk as a function of the values of the two decision variables. We tested the hypotheses that observers base their 2AFC decisions on (a) the difference between the two decision variables, (b) independent yes-no decisions on the two decision variables, or (c) just one of the decision variables. We found that all observers' decision spaces had a triangular guessing region, which is not predicted by any of the above models. However, this finding is consistent with model (a) plus intrinsic uncertainty. We conclude that the classic difference model favoured by detection theory is a valid model of 2AFC decisions. In Experiment 2, four observers discriminated between black and white Gaussian disks at fixation, and the two stimulus intervals were separated in time (1000 ms) rather than space. Again observers' decision spaces supported the difference model. We discuss how proxy decision variables can be used to test a wide range of additional signal detection models in domains such as cue combination and visual search. Meeting abstract presented at VSS 2014
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.023 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".