Get over time: a longitudinal variationist analysis of passive voice in contemporary English
Bibliographic record
Abstract
The English voice system has two passive auxiliaries: the canonical be-passive, and the more recent get-passive. Accounts of the get-passive in the linguistic literature draw from descriptive, historical, corpus linguistic, and variationist perspectives. Much existing work on the get-passive from the former three traditions notes semantic dissimilarities from the be-passive, suggesting that these two forms are not interchangeable and therefore do not constitute a typical sociolinguistic variable. Nonetheless, variationist work has treated the be- and get-passives as alternants expressing the same function. This latter work has focused on social factors alone, setting aside purported linguistic differences. This thesis provides a variationist account of the be- and get-passives, considering not only social factors, but also operationalizing as linguistic factors previously noted semantic characteristics, demonstrating which factors constrain variation and providing a holistic picture of the get-passive in vernacular English. The speakers in this study span a birth range of 1865 to 1996, providing a longitudinal scope from which to view the grammaticalization of the feature. Following the principle of accountability (Labov, 1972), instances of be- and get-passives were extracted from 108 speakers born and raised in Victoria, British Columbia, Canada (N=1716). Distributional and inferential results show a substantial increase in rates of get-passive over the last 130 years, indicating an active and ongoing change in progress. Social and linguistic factors alike are shown to play meaningful roles in variant selection, revealing a (largely) longitudinally stable variable grammar. The longitudinal scope of the study illuminates grammaticalization pathways into the 20th century and reinforces attested semantic links between the contemporary get-passive and its proposed lexical source(s).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.000 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".