Disentangling Genuine Semantic Stroop Effects in Reading from Contingency Effects: On the Need for Two Neutral Baselines
Bibliographic record
Abstract
The automaticity of reading is often explored through the Stroop effect, whereby color-naming is affected by color words. Color associates (e.g., "sky") also produce a Stroop effect, suggesting that automatic reading occurs through to the level of semantics, even when reading sub-lexically (e.g., the pseudohomophone "skigh"). However, several previous experiments have confounded congruency with contingency learning, whereby faster responding occurs for more frequent stimuli. Contingency effects reflect a higher frequency-pairing of the word with a font color in the congruent condition than in the incongruent condition due to the limited set of congruent pairings. To determine the extent to which the Stroop effect can be attributed to contingency learning of font colors paired with lexical (word-level) and sub-lexical (phonetically decoded) letter strings, as well as assess facilitation and interference relative to contingency effects, we developed two neutral baselines: each one matched on pair-frequency for congruent and incongruent color words. In Experiments 1 and 3, color words (e.g., "blue") and their pseudohomophones (e.g., "bloo") produced significant facilitation and interference relative to neutral baselines, regardless of whether the onset (i.e., first phoneme) was matched to the color words. Color associates (e.g., "ocean") and their pseudohomophones (e.g., "oshin"), however, showed no significant facilitation or interference relative to onset matched neutral baselines (Experiment 2). When onsets were unmatched, color associate words produced consistent facilitation on RT (e.g., "ocean" vs. "dozen"), but pseudohomophones (e.g., "oshin" vs. "duhzen") failed to produce facilitation or interference. Our findings suggest that the Stroop effects for color and associated stimuli are sensitive to the type of neutral baseline used, as well as stimulus type (word vs. pseudohomophone). In general, contingency learning plays a large role when repeating congruent items more than incongruent items, but appropriate pair-frequency matched neutral baselines allow for the assessment of genuine facilitation and interference. Using such baselines, we found reading processes proceed to a semantic level for familiar words, but not pseudohomophones (i.e., phonetic decoding). Such assessment is critical for separating the effects of genuine congruency from contingency during automatic word reading in the Stroop task, and when used with color associates, isolates the semantic contribution.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.015 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".