Bibliographic record
Abstract
One of the most notable demolinguistic phenomena of the modern age has been the expansion of the English language, from its roots as a set of West Germanic dialects in early medieval England to its current position as the leading global lingua franca. It now has hundreds of millions of native speakers and an even larger population of non-native speakers, living in every region of the world. This expansion has involved three major phases: the anglicization of Britain's Celtic population; the transfer of English to other continents through emigration from Britain and colonialism; and the adoption of English as an international language by people in non-English-speaking countries beyond the former British Empire (the three diasporas of Kachru, Kachru and Nelson 2006, originally conceived by Kachru 1985). Part of the middle phase of expansion, beginning in the seventeenth century, was the bringing of English to North America by British colonists. These were as much Irish and Scottish as they were English, thereby reflecting the initial phase of expansion. The eventual success of their colonial project drew many more settlers, first from Britain and Europe and then from all over the world. If they did not already speak English, most of these settlers adopted it and most of their children became native speakers, so that English was established as the majority language of two new multi-ethnic nations, the United States and Canada. This book is a study of the English language in Canada: its current status, history and most important characteristics.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.006 |
| Science and technology studies | 0.029 | 0.007 |
| Scholarly communication | 0.010 | 0.003 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.019 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".