Progressing, not regressing: a possible solution to the problem of regression to the mean in unconscious processing studies
Bibliographic record
Abstract
How convincing is current evidence for unconscious processing? Recently, a major criticism suggested that some, if not much, of this evidence might be explained by a mere statistical phenomenon: regression to the mean (RttM). Excluding participants based on an awareness assessment is a common practice in studies of unconscious processing, and this post-hoc data selection might lead to false effects that are driven by RttM for aware participants wrongfully classified as unaware. Here, we examined this criticism using both simulations and data from 12 studies probing unconscious processing (35 effects overall). In line with the original criticism, we confirmed that the reliability of awareness measures in the field is concerningly low. Yet using simulations, we showed that reliability measures might be unsuitable for estimating error in awareness measures. Furthermore, we examined other solutions for assessing whether an effect is genuine or reflects RttM; all suffered from substantial limitations, such as a lack of specificity to unconscious processing, lack of power, or unjustified assumptions. Accordingly, we suggest a new nonparametric solution, which enjoys high specificity and relatively high power. Together, this work emphasizes the need to account for measurement error in awareness measures and evaluate its consequences for unconscious processing effects. It further suggests a way to meet the important challenge posed by RttM, in an attempt to establish a reliable and robust corpus of knowledge in studying unconscious processing.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".