MétaCan
Menu
Back to cohort
Record W7081981260 · doi:10.17605/osf.io/u3qx7

Understanding Prejudice in generative AI Through the Lens of Social Psychology: A Systematic Review

2025· other· en· W7081981260 on OpenAlexaboutno aff

Bibliographic record

VenueOpen Science Framework · 2025
Typeother
Languageen
FieldMedicine
TopicPrenatal Screening and Diagnostics
Canadian institutionsnot available
Fundersnot available
KeywordsGenerative grammarPrejudice (legal term)Field (mathematics)Social learningSocial cognitionIntersection (aeronautics)

Abstract

fetched live from OpenAlex

REVIEW TEAM MEMBERS Elena Trifiletti (review contact) ORCID: 0000-0002-9203-507x. Università di Verona. Italy, Associate Professor. Nuccio Ludovico. ORCID: 0000-0003-3640-9775. Università di Modena e Reggio Emilia. Italy, Post-Doc Fellow. Jessica Boin. ORCID: 0000-0001-6141-1274. Università di Padova. Italy, Assistant Professor. Loris Vezzali. ORCID: 0000-0001-7536-9994. Università di Modena e Reggio Emilia. Italy, Full Professor. COUNTRY Italy KEYWORDS Prejudice; Stereotype; Bias; Discrimination; Generative Artificial Intelligence; Social Psychology theories RATIONALE Generative artificial intelligence (AI) systems have become increasingly integrated into everyday life. Concerns about the reproduction and amplification of human biases - especially prejudice and stereotypes related to gender, ethnicity and race, sexual orientation, and other social categories – are growing. While computer scientists are increasingly working to identify and reduce bias in AI systems, it is critical to examine these phenomena through the lens of social psychology. Social psychology offers robust theories on prejudice, stereotyping, and intergroup processes and the research in this field provides valuable insights into how biases form, persist, and manifest in human behavior. These insights that can help explain and predict how such biases are encoded and reproduced by generative AI models. However, insights from generative AI and social psychology remain largely disconnected, and integrating knowledge from both fields would be highly beneficial. A systematic review is needed to synthesize existing literature at the intersection of social psychology and AI research and to offer a framework for interpreting social bias in generative AI. This review can guide both social psychology research and the ethical development of AI technologies, fostering interdisciplinary collaboration and socially responsible innovation. REVIEW OBJECTIVES This review has two main research questions: (1) What social psychological theories, models, or concepts are being used to understand and investigate social bias in generative AI systems? (2) How they can help explain the ways prejudice is encoded or perceived in generative AI? INCLUSION CRITERIA (type of studies) This systematic review will include: - empirical studies employing Machine Learning, Deep Learning, Natural Language Processing (NLP), Computer Vision, Robotics, or Expert Systems - theoretical studies; - studies must address social bias (i.e., prejudice, stereotyping, or discrimination) in AI-generated content; - studies must draw on one or more social psychology theories or concepts when addressing social bias. INCLUSION CRITERIA (human participants) The review will not primarily focus on studies involving human participants. However, if studies meeting the inclusion criteria do include human participants, they will be eligible regardless of age group or user type, including general users, consumers, or professionals. INCLUSION CRITERIA (publication type) Articles; Conference proceedings INCLUSION CRITERIA (publication status) Published; early access; pre-print. Unpublished studies will not be sought. CONTEXT This systematic review will include studies conducted in any context or setting, including but not limited to computer science laboratories and applied domains, without restrictions on domain, population, or application area. TIMELINE OF THE REVIEW Start date: 27 August 2025. End date: 31 January 2026. DATE OF REGISTRATION IN OSF 26 August 2025 DATABASES THAT WILL BE SEARCHED Scopus; Web of Science Core Collection; IEEE Xplore; ACM digital library. SEARCH LANGUAGE DESCRIPTION The review will only include studies published in English. SEARCH DATE RESTRICTIONS Databases will be searched for articles published from 2022 (i.e., following the public release of ChatGPT, which significantly increased access to large language models). SEARCH STRATEGY The selected libraries and repositories will be systematically explored using advanced search queries. The query has been designed to capture the intersection of two core concepts relevant to this review: on the one hand, generative AI solutions and models; on the other hand, social bias and stereotypes. This approach led to the construction of two distinct keyword blocks: • Block I: (“large language model*" OR "llm*" OR "transformer-based model*" OR "pre-trained language model*" OR "pre-trained transformer*" OR "generative ai" OR "generative artificial intelligence" OR "gpt" OR "foundation model*" OR "vision foundation model*" OR "large vision model*" OR “vision language model*” OR "captioning model*" OR "chatgpt" OR "claude ai" OR "copilot" OR "gemini" OR "llama" OR "deepseek" OR "grok" OR "midjourney" OR "dall-e 2" OR "dall-e 3" OR “clip model” OR "stable diffusion" OR "adobe firefly" OR "nightcafe" ) • Block II: ( "prejudice*" OR "stereotype*" OR "racism" OR "sexism" OR "ageism" OR "genderism" OR "ableism" OR "ethnocentrism" OR "social stigma" OR "dehumanization" OR "ethnicity" OR "outgroup attitude*" OR "intergroup attitude*" OR "implicit bias*" OR "explicit bias*" OR "intergroup bias*" OR "ingroup bias*" OR "racial bias*" OR "ethnic* bias*" OR "gender bias*" OR "sociodemographic bias*" OR "age bias*" OR "disability bias*" OR "attitude bias*" OR "social discrimination*" OR "racial discrimination*" OR "ethnic* discrimination*" OR "gender discrimination*" OR "sociodemographic discrimination*" OR "age discrimination*" OR "disability discrimination*" OR "group-based discrimination*" OR "attitude discrimination*" ) The “*” wildcard is used to include term variants where needed. An “AND” operator is placed between the two keyword blocks. The query will be applied to titles, abstracts, and, where available, metadata keywords of each record. Inclusion criteria and search restrictions will either be embedded directly into the query or subsequently applied to the retrieved records through additional filtering steps. SELECTION PROCESS Studies will be screened independently by at least two people (or person/machine combination) with a process to resolve differences. DATA EXTRACTION Data will be extracted independently by at least two people (or person/machine combination) with a process to resolve differences. Authors will be asked to provide any required data not available in published reports. STUDY RISK OF BIAS OR QUALITY ASSESSMENT Risk of bias will be assessed using: - Cochrane RoB-2: J.A.C. Sterne, J. Savović, M.J. Page, R.G. Elbers, N.S. Blencowe, I. Boutron, C.J. Cates, H.Y. Cheng, M.S. Corbett, S.M. Eldridge, J.R. Emberson, M.A. Hernán, S. Hopewell, A. Hróbjartsson, D.R. Junqueira, P. Jüni, J.J. Kirkham, T. Lasserson, T. Li, …, J.P.T. Higgins (2019). RoB 2: A revised tool for assessing risk of bias in randomised trials. BMJ, 366, p. l4898. https://doi.org/10.1136/bmj.l4898. - Kmet, L. M., Lee, R. C., & Cook, L. S. (2004). Standard quality assessment criteria for evaluating primary research papers from a variety of fields. Alberta Heritage Foundation for Medical Research. https://doi.org/10.7939/R37M04F16 Data will be assessed independently by at least two people (or person/machine combination) with a process to resolve differences. Additional information will be sought from study investigators if required information is unclear or unavailable in the study publications/reports. OUTCOMES TO BE ANALYZED 1. Social psychological theories and concepts applied in AI bias studies Description: identification and categorization of social psychological theories and concepts (e.g., social identity theory, implicit/explicit bias) that are explicitly or implicitly referenced in studies addressing AI bias. Rationale: to map the theoretical underpinnings informing AI bias research and interventions. 2. Operationalization of social psychological concepts in methodology or analysis Description: examining how psychological theories/concepts are operationalized (e.g., in designing AI models, framing experimental hypotheses, interpreting algorithmic outcomes). Rationale: to assess how deeply or superficially these theories are integrated into the study designs. 3. Strategies and interventions for bias mitigation that are informed by social psychology Description: description and classification of mitigation strategies (e.g., debiasing techniques, fairness-aware algorithms) that are grounded in social psychological principles. Rationale: To evaluate the translation of theory into practice in addressing bias in AI systems. 4. Effects of theory-informed bias mitigation strategies/interventions Description: quantitative (e.g., changes in algorithmic fairness metrics) and qualitative (e.g., user perceptions or trust) outcomes of mitigation strategies. Rationale: To assess the reported effectiveness and impact of interventions derived from social psychological theories/concepts. 5. Types of bias and stereotypes addressed Description: to identify and classify the kinds of bias/stereotypes targeted in the AI studies (e.g., gender bias, racial/ethnic bias, algorithmic discrimination, representational harm). Rationale: to analyze which forms of social bias receive more attention or are neglected in AI research using social psychological concepts. 6. AI Model(s) used Description: to identify and classify the types of AI models employed in the studies (e.g., large language models, large visual models, multimodal models). Rationale: to examine which categories of generative AI models are most frequently studied in relation to bias and stereotypes by using social psychological concepts, and to assess whether specific models/architectures are particularly prone to bias/stereotype occurrence. STRATEGY FOR DATA SYNTESIS The expectation is that the synthesis will be narrative and qualitative and will not engage in a meta-analysis. The software EndNote will be used to extract studies' characteristics and outcome data. Two reviewers will extract the d

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.005
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Systematic review · Consensus signal: none
GenreCandidate signal: Other · Consensus signal: Other
Teacher disagreement score0.893
Threshold uncertainty score0.580

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.005
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.002
Science and technology studies0.0000.001
Scholarly communication0.0000.000
Open science0.0010.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.143
GPT teacher head0.440
Teacher spread0.298 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSystematic review
Domainnot available
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueOpen Science FrameworkSame topicPrenatal Screening and DiagnosticsFrench-language works237,207