Gender Markedness: A Corpus-Based Analysis of the Epicene Pronouns ‘S/He’ and ‘He/She’
Bibliographic record
Abstract
Gender reference emerges in the speaker's positionality when addressing interlocutors in discourse and is perceived and interpreted through cultural, social, and linguistic lenses. Based on a semantic system, the English language denotes gender with grammatically specific pronouns and lexical unmarked vs. marked forms, mainly concerning job-related titles or honorifics. As part of the unmarked category, the marked lexemes represent the variant to the norm: they are formally larger, and depend on the context, requiring extra-linguistic effort, either in production or comprehension (Givón, 1995, pp. 25-28). In this sense, the occurrence of epicene pronouns such as s/he, compared to he/she in the British Web Corpus (ukWaC), sheds light on how text interpretation in context may promote gender inclusivity through more equitable linguistic practices within different contexts. A corpus-based approach thus investigated these referents in their concordances to examine their usage and implications of meaning to understand the textual genre wherein they are preferably used. Two research questions arise: 1) What are the frequency distributions and collocational patterns of the gender-neutral pronoun “s/he” and unmarked/marked gendered pronoun “he/she” in the ukWaC corpus? 2) What semantic and grammatical roles do the collocates of “s/he” and “he/she” play, and how do these roles differ between the two pronouns? This study aims to provide insights to assess how these forms operate across genres and contexts, focusing on their functional, pragmatic, and institutional roles versus their stylistic implications.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.065 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".