MétaCan
Menu
Back to cohort
Record W7112916737

Can we interpret gender? Using contrastive explanations to understand gender choices by translation systems

2025· article· en· W7112916737 on OpenAlexaboutno aff

Bibliographic record

VenueGhent University Academic Bibliography (Ghent University) · 2025
Typearticle
Languageen
FieldComputer Science
TopicExplainable Artificial Intelligence (XAI)
Canadian institutionsnot available
Fundersnot available
KeywordsInterpretabilitySentenceFocus (optics)Context (archaeology)Affect (linguistics)PerceptionVocabularyInterpretation (philosophy)Relevance (law)
DOInot available

Abstract

fetched live from OpenAlex

Interpretability can be implemented as a means to understand decisions taken by (black box) models, such as machine translation (MT) systems or large language models (LLMs). Yet, this has barely been explored in relation to a long-lasting problem in these systems: gender bias, which, although widely studied, has not yet been solved (Savoldi et al., 2025). One study by Attanasio et al. (2023) has explored the interplay between these two domains and found that interpretability is a “valuable tool for studying and mitigating bias in language models”. This interplay is the focus of this research. It has been shown that certain contextual cues influence a human’s perception of gender in an ambiguous sentence context (Hackenbuchner et al., forthcoming). We aim to understand whether the same contextual cues influence an MT system when it translates a person’s gender into a certain inflection. To study this, we look at saliency scores to analyse which input tokens are most relevant for a certain translation decision. Specifically, we focus on contrastive explanations, which have been shown to outperform non-contrastive ones, to analyse which input tokens lead a model to produce one output instead of another (why is X predicted instead of Y?), as introduced in Yin and Neubig (2022). This interpretability technique helps understand which contextual cues (input tokens) in the source sentence affect the model’s choice of a certain gender inflection instead of another in the target translation. This research is conducted on the dataset of gender-ambiguous sentences introduced in Hackenbuchner et al. (forthcoming) and compared to that study’s human annotations of contextual cues affecting their gender perceptions. Using the inseq toolkit (Sarti 2023), we realise contrastive explanations by computing the difference between two attribution outputs (the difference in probability between two options) based on the target translation (taken from the original dataset) and a translation contrasting in terms of gender. To exemplify this, we analyse which input tokens in the source “The business writer from Miami” lead to a higher probability of being translated into male (e.g., DE: Schriftsteller) or into female (e.g., DE: Schriftstellerin). Preliminary results show that there is a noticeable overlap between human perceptions and model attribution. Contextual cues (words) that influence human perception of gender are among the most frequent tokens that influence a model’s translation of gender in the target (i.e. that increase the probability of a certain gender output over another). This shows that humans and models seem to be influenced by very similar (if not the same) contexts in regards to gender (even in ambiguous scenarios). With this study, we contribute to the very limited research conducted on interpretability measures of model decisions in the translation of gender. ---------------------------------------------------------------------------------------------------------------------------------------------- Beatrice Savoldi, Jasmijn Bastings, Luisa Bentivogli, Eva Vanmassenhove. 2025. A decade of gender bias in machine translation. In: Patterns. Volume 6, Issue 6. Gabriele Sarti, Nils Feldhus, Ludwig Sickert, and Oskar van der Wal. 2023. Inseq: An Interpretability Toolkit for Sequence Generation Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pages 421–435, Toronto, Canada. Association for Computational Linguistics. Giuseppe Attanasio, Flor Miriam Plaza del Arco, Debora Nozza, and Anne Lauscher. 2023. A Tale of Pronouns: Interpretability Informs Gender Bias Mitigation for Fairer Instruction-Tuned Machine Translation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 3996–4014, Singapore. Association for Computational Linguistics. Janiça Hackenbuchner, Arda Tezcan, Joke Daems. Forthcoming: CLIN vol 14. Gender Bias and the Role of Context in Human Perception and Machine Translation. Kayo Yin and Graham Neubig. 2022. Interpreting Language Models with Contrastive Explanations. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 184–198, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Bibliometrics
Consensus categoriesBibliometrics
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.968
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0230.032
Science and technology studies0.0010.000
Scholarly communication0.0000.002
Open science0.0020.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.065
GPT teacher head0.270
Teacher spread0.205 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designTheoretical or conceptual
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueGhent University Academic Bibliography (Ghent University)Same topicExplainable Artificial Intelligence (XAI)French-language works237,207