Commentary: Cognitive reflection vs. calculation in decision making
Bibliographic record
Abstract
Sinayev and Peters (2015; hereafter S&P) present two competing hypotheses to explain performance on the Cognitive Reflection Test (CRT).They dub the first the "Cognitive Reflection Hypothesis" and attribute it to other researchers: "Each of these researchers assumes that differences in CRT performance indicated differences in the ability to detect and correct incorrect intuitions. . ." and ". . .implicitly assume that numerical ability is an irrelevant detail when it comes to solving CRT and related problems" (p.2).They contrast this with their "Numeracy Hypothesis" which states that "the CRT is primarily a measure of numeric ability" (p.3).S&P report two studies whose results, they argue, favor the Numeracy Hypothesis over the Cognitive Reflection Hypothesis.They conclude that numeric ability is "the key mechanism" that explains the association between CRT performance and decision making (p.1), although they also state that the ability to detect and correction intuitions (apart from numeracy) plays a role in CRT performance.Both of the hypotheses presented by S&P emphasize the role of cognitive ability in CRT performance.In this commentary we introduce an alternative hypothesis that was not discussed by S&P; namely, that the propensity or disposition to think analytically plays an important role in CRT performance (Pennycook et al., 2015b).We discuss recent empirical evidence that supports the claim that the CRT is more than just a measure of numeracy or, more generally, cognitive ability. DISTINGUISHING COGNITIVE ABILITY AND ANALYTIC COGNITIVE STYLEDual process theorists often distinguish between disposition and ability as factors that determine good reasoning (e.g., Stanovich and West, 2000;Stanovich, 2009;Evans and Stanovich, 2013).The logic is as follows: If someone does not have the disposition or willingness to think analytically, they will not fully exercise their cognitive ability and will not do as well on the problem.Naturally, the converse is also true: If someone does not have sufficient cognitive ability, it will not matter how much time and effort they are willing to spend thinking about the problem.This distinction has been applied to CRT performance.For example, according to Toplak et al. (2014): "the CRT is a measure of the tendency toward the class of reasoning error that derives from miserly processing.This may be why the predictive power of the CRT is in part separable from cognitive ability.The latter measures computational power that is available to the individual, but not necessarily the depth of processing that is typically used in most situations" (p.165).That
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.056 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.005 | 0.011 |
| Scholarly communication | 0.004 | 0.007 |
| Open science | 0.009 | 0.003 |
| Research integrity | 0.050 | 0.057 |
| Insufficient payload (model declined to judge) | 0.007 | 0.008 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".