The connection between heritage and endangered languages
Bibliographic record
Abstract
In this presentation, I investigate the link between heritage (immigrant) languages and indigenous endangered languages. A heritage language (HL) is a minority language learned in the home by speakers who are more dominant in the majority societal language. An endangered language (EL) is a language that is at risk of falling out of use, due to the scarcity of surviving speakers and lack of intergenerational transmission. Connections between the two types of languages, both minoritized, have not been investigated in a systematic and extensive way. Aside from letting us understand social and cultural pressures associated with language shift, focusing on structural parallels between HLs and ELs will allow us to conduct more inclusive research on minoritized languages.\nIndigenous ELs, languages that are not robustly transmitted to younger generations, share important characteristics with immigrant HLs (Sasse 1992): (i) the switch from early and naturalistic immersion in the ancestral language to takeover by the ambient language, e.g., in the context of Residential Schools in Canada or the USA, and (ii) the presence of socio-economic power associated with the ambient language. In both immigrant HL and indigenous EL settings, this socio-cultural dynamic gives rise to a range of bilingual outcomes. However, unlike HLs, there is no baseline because the traditional language is lost. Thus, identifying structural properties that arise due to extensive bilingualism leads to a better analysis of the current state of ELs. This is where comparisons to HLs are particularly fruitful and effective.\nThe gain for linguistic theory in connecting ELs and HLs is twofold. First, examining languages that are used in the context of extreme bilingualism would allow us to better understand the nature of universal structural principles; second, some unusual phenomena that may be observed in ELs could be explained by effects of recessive bilingualism, which in turn would prevent the unnecessary exotification of such languages. In my talk, I will present and analyze particular structural parallels, which show that both types of languages observe locality, maximize the use of anaphoric dependencies, and show a bias against using displacement as a structure building mechanism. I propose main mechanisms that influence the grammatical structure of ELs and HLs and present examples of structural changes due to the operation of these mechanisms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.002 | 0.006 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.001 | 0.005 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".