An efficient and reliable DNA-based sex identification method for archaeological Pacific salmonid (Oncorhynchus spp.) remains
Bibliographic record
Abstract
Pacific salmonid (Oncorhynchus spp.) remains are routinely recovered from archaeological sites in northwestern North America but typically lack sexually dimorphic features, precluding the sex identification of these remains through morphological approaches. Consequently, little is known about the deep history of the sex-selective salmonid fishing strategies practiced by some of the region's Indigenous peoples. Here, we present a DNA-based method for the sex identification of archaeological Pacific salmonid remains that integrates two PCR assays that each co-amplify fragments of the sexually dimorphic on the Y chromosome (sdY) gene and an internal positive control (Clock1a or D-loop). The first assay co-amplifies a 95 bp fragment of sdY and a 108 bp fragment of the autosomal Clock1a gene, whereas the second assay co-amplifies the same sdY fragment and a 249 bp fragment of the mitochondrial D-loop region. This method's reliability, sensitivity, and efficiency, were evaluated by applying it to 72 modern Pacific salmonids from five species and 75 archaeological remains from six Pacific salmonids. The sex identities assigned to each of the modern samples were concordant with their known phenotypic sex, highlighting the method's reliability. Applications of the method to dilutions of modern DNA samples indicate it can correctly identify the sex of samples with as little as ~39 pg of total genomic DNA. The successful sex identification of 70 of the 75 (93%) archaeological samples further demonstrates the method's sensitivity. The method's reliance on two co-amplifications that preferentially amplify sdY helps validate the sex identities assigned to samples and reduce erroneous identifications caused by allelic dropout and contamination. Furthermore, by sequencing the D-loop fragment used as a positive control, species-level and sex identifications can be simultaneously assigned to samples. Overall, our results indicate the DNA-based method reported in this study is a sensitive and reliable sex identification method for ancient salmonid remains.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".