Contact-induced lexical development in Yupik and Inuit languages
Bibliographic record
Abstract
Lexical change in Yupik and Inuit languages was relatively slow until the period of widespread cultural change brought about by contact with Europeans over the past few centuries, probably because there had been little earlier contact with other language families. The colonial period brought various groups to the Arctic and different waves of language contact, primarily with Danish, French, English, and Russian. Lexical borrowing has been significant, and old borrowings, often the result of early trade, can be distinguished from later ones and often pertain to food, tobacco, tools, fabric and other areas where new goods were introduced. Later borrowings came about largely when European political structures were set up and may be less thoroughly integrated phonologically than older borrowings. Numbers of borrowings can be taken to reflect the extent of the foreign contact, as is clearly the case with Russian words in Alaskan languages, most numerous in Aleut, which had the most sustained Russian presence. New religious terms to describe Christianity came into the languages during the colonial period, sometimes as borrowings, but also as relexicalizations of old words pertaining to shamanism. A third means of lexical expansion is coinage, where new terms are invented based on native roots and suffixes. The languages and dialects may thus develop words for the same object or concept by borrowing from different European languages, by relexicalizing an old word, or by coining a new one, with a different result in each case. Different sources for new lexical items have resulted in an important level of differentiation among the languages, and this differentiation needs to be recognized.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.009 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".