Bibliographic record
Abstract
Introduction David Crystal estimates that about 400 million people have English as their first language, and that in total as many as 1500 million may be to a greater or lesser extent fluent speakers of English (see Chapter 9, Table 9.1). The two largest countries (in terms of population) where English is the inherited national language are Britain and the USA. But it is also the majority language of Australia and New Zealand, and a national language in both Canada and South Africa. Furthermore, in other countries it is a second language, in others an official language or the language of business. If, more parochially, we restrict ourselves to Britain and the USA, the fact that it is the inherited national language of both does not allow us to conclude that English shows a straightforward evolution from its ultimate origins. Yet originally English was imported into Britain, as also happened later in North America. And in both cases the existing languages, whether Celtic, as in Britain, or Amerindian languages, as in North America, were quickly swamped by English. But in both Britain and the USA, English was much altered by waves of immigration. Chapter 8 will demonstrate how that occurred in the USA. In Britain, of course, the Germanic-speaking Anglo-Saxons brought their language with them as immigrants. The eighth and ninth centuries saw Scandinavian settlements and then the Norman Conquest saw significant numbers of French-speaking settlers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.228 | 0.137 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".