Not All Stroop-Type Tasks Are Alike: Assessing the Impact of Stimulus Material, Task Design, and Cognitive Demand via Meta-analyses Across Neuroimaging Studies
Bibliographic record
Abstract
The Stroop effect is one of the most often studied examples of cognitive conflict processing. Over time, many variants of the classic Stroop task were used, including versions with different stimulus material, control conditions, presentation design, and combinations with additional cognitive demands. The neural and behavioral impact of this experimental variety, however, has never been systematically assessed. We used activation likelihood meta-analysis to summarize neuroimaging findings with Stroop-type tasks and to investigate whether involvement of the multiple-demand network (anterior insula, lateral frontal cortex, intraparietal sulcus, superior/inferior parietal lobules, midcingulate cortex, and pre-supplementary motor area) can be attributed to resolving some higher-order conflict that all of the tasks have in common, or if aspects that vary between task versions lead to specialization within this network. Across 133 neuroimaging experiments, incongruence processing in the color-word Stroop variant consistently recruited regions of the multiple-demand network, with modulation of spatial convergence by task variants. In addition, the neural patterns related to solving Stroop-like interference differed between versions of the task that use different stimulus material, with the only overlap between color-word, emotional picture-word, and other types of stimulus material in the posterior medial frontal cortex and right anterior insula. Follow-up analyses on behavior reported in these studies (in total 164 effect sizes) revealed only little impact of task variations on the mean effect size of reaction time. These results suggest qualitative processing differences among the family of Stroop variants, despite similar task difficulty levels, and should carefully be considered when planning or interpreting Stroop-type neuroimaging experiments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.046 | 0.096 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.008 | 0.029 |
| Bibliometrics | 0.007 | 0.008 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.005 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".