Bibliographic record
Abstract
The Languages of the PacificWhen different people speak of the Pacific region, they often mean different things.In some senses, people from such Pacific Rim countries as Japan and Korea, Canada and the United States, and Colombia and Peru are as much a part of the region as are those from Papua New Guinea, Fiji, the Marshall Islands, Tonga, and so on.In this book, however, I use the term "the Pacific" to refer to the island countries and territories of the Pacific Basin, including Australia and New Zealand.This Pacific has traditionally been divided into four regions: Melanesia, Micronesia, Polynesia, and Australia (see map 2).Australia is clearly separate from the remainder of the Pacific culturally, ethnically, and linguistically.The other three regions are just as clearly not separate from one another according to all of these criteria.There is considerable ethnic, cultural, and linguistic diversity within each of these regions, and the boundaries usually drawn between them do not necessarily coincide with clear physical, cultural, or linguistic differences.These regions, and the boundaries drawn between them, are largely artifacts of the western propensity, even weakness, for classification, as the continuing and quite futile debate over whether Fijians are Polynesians or Melanesians illustrates.Having said this, however, I will nevertheless continue to use the terms "Melanesia," "Micronesia," and "Polynesia" to refer to different geographical areas within the Pacific basin, without prejudice to the relationships of the languages or the cultures of people of each region. How Many Languages?This book deals mainly with the indigenous languages of the Pacific region.There are many other languages that can be called "Pacific languages," for
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.004 | 0.002 |
| Scholarly communication | 0.006 | 0.004 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.004 |
| Insufficient payload (model declined to judge) | 0.081 | 0.016 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".