When and Why are Social Categories Overused Relative to Individuating Information? A Bayesian Approach to Identifying Biases in Impression Formation Processes
Bibliographic record
Abstract
When making inferences about others, people often need to integrate multiple types of information, including information about the person’s social categories (e.g., occupation) as well as other ‘individuating’ information (e.g., their behaviors). The current work re-examines this integration process, to understand when and why it might lead to biases that involve over-relying on category information. To do so, we identify two key challenges in identifying such biases, and develop a novel Bayesian modelling approach to overcome these challenges. As a first step in applying this approach, the current work examined a set of novel predictions based on viewing the Continuum Model of impression formation in light of this Bayesian approach. Specifically, a series of six studies tested whether occupation categories might be overused in general, or especially in conditions thought to reduce effortful processing (i.e., greater cognitive load or information consistency). At baseline, there was no consistent evidence of category overuse in any of these conditions, speaking against the idea that intrinsic differences in how these categories are processed or represented will lead to their overuse. In addition, this work provided the first direct evidence of a case where categories were overused: when contextual factors (i.e., background goals) made them the category especially relevant, while processing resources were limited. More broadly, the current work developed the theoretical and methodological foundations for identifying category overuse that stems from biased inference processes, and demonstrated the power of this approach for understanding when and why these biases occur.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".