Bibliographic record
Abstract
For the many communities where knowledge of the indigenous language has declined over the past century, early vocabularies of the language and longer archival and published linguistic works can serve as invaluable tools for retrieving forgotten or lost words when preparing a dictionary. As a language declines in use, there are ever fewer opportunities for language learners to hear the language. As a result, they may not learn certain terms that occur only rarely and may substitute more frequent terms when the need arises. For example, in the past, the Tuscarora language possessed several different words to name different kinds of feathers, including uhraOneh large feather, wing feather, quill, uhsnuOkreh small or body feather, uθnuureh feather, down; and yuhraOkwaOr tail feather. Today, only the word uhraOneh is in common use in the eastern dialect of the language spoken on the Tuscarora Indian Reservation in New York State and only the word uhsnuOsreh (from earlier uhsnuOkreh) was recorded from the last speakers of the western dialect of the language on the Six Nations Reserve in Ontario. In addition, individuals who spoke a language fluently as a child, but have not used the language since, often times repress their knowledge of the language and need some external stimulus to jog their memory. Also, all languages change over time. Some words are replaced by new words, while other words that have outlived their usefulness are lost from the language. Early vocabularies of a language can help a community in these situations to retrieve lost or forgotten vocabulary.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".