Creating the Blackfoot digital library: the challenge of cultural sensitivity.
Bibliographic record
Abstract
In the mid 1990’s it was estimated that there are only about 5,000 – 8,000 speakers of the \nBlackfoot language and that the numbers were declining. The decline of the language was \nexacerbated by the absence of a generally accepted writing system. The orthography most \ncommonly used for writing Blackfoot on the three Southern Alberta reserves was only approved as the official writing system in 1975. \nThis resulted in very little written material being produced by the Blackfoot people that captured their history In 2006 the University of Lethbridge and Red Crow Community College joined forces in to ensure \nthat as much as possible of the Blackfoot cultural record will be preserved and made accessible \nthrough the creation of a Blackfoot Digital Library. \nA foundational requirement of the digital library was cultural sensitivity and specifically that it must appropriately honor the Blackfoot worldview. In the traditional Blackfoot worldview the underlying premise is that all knowledge is derived from place, which posed a significant challenge. The solution was the design of a custom-made search interface that displays the search results on a \ndigital map where the user is immediately confronted with the land. The map, displays the \nplace(s) where the assets that were retrieved by the search, originated from and offer access to the various assets themselves. \nThe presentation informs on the challenges faced by the initiative.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.020 | 0.031 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.019 | 0.018 |
| Scholarly communication | 0.040 | 0.026 |
| Open science | 0.004 | 0.032 |
| Research integrity | 0.005 | 0.006 |
| Insufficient payload (model declined to judge) | 0.015 | 0.008 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".