Boreal Bias: A Critique of CRM Testing Methodologies in Alberta’s Boreal Forest
Bibliographic record
Abstract
This project, conducted in partnership with Ermineskin Cree Nation through the Ermineskin Industrial Relations Department (EIRD), was undertaken to address biases present in Alberta Culture Resource Management (CRM) archaeology that skew the data and perpetuate the conception that boreal archaeological sites are generally small and ephemeral. A comprehensive analysis of archaeological sites in the study area identified site clusters – where two or more archaeological sites have been identified within 100m of each other – to determine whether these previously identified sites are isolated with clearly defined boundaries, or if they could be connected as part of a larger site area. To this end, subsurface testing was conducted on the untested terrain between the previously identified site boundaries (FgPw-37, FgPw-41, and FgPw-43) within site cluster FID1127. The site cluster is located on a large parabolic sand dune near the contemporary Tidewater Gas Plant in the foothills of west-central Alberta. The testing was designed to ascertain whether current CRM methodologies adequately identify cultural material and accurately reflect the special extent of known sites. Identification of cultural material between these known sites supports the theory that current CRM practices in the Alberta boreal forest are missing key archaeological material, and as a result, breaking up larger habitation areas into what appear to be small ephemeral sites. This is not an accurate reflection of past lifeways in this region. Because industrial expansion in the boreal forest shows no evidence of slowing down and archaeological material is not a renewable resource, CRM methodologies and regulations must be continually tested and updated. Failure to do so compromises archaeological resources and contributes to the erasure of Indigenous history in Alberta.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.288 | 0.365 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.007 | 0.010 |
| Science and technology studies | 0.006 | 0.017 |
| Scholarly communication | 0.009 | 0.004 |
| Open science | 0.011 | 0.006 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".