Genomic sex identification of ancient pinnipeds using the dog genome
Bibliographic record
Abstract
Determining the proportion of males and females in zooarchaeological assemblages can be used to reconstruct the diversity and severity of past anthropogenic impacts on animal populations, and can also provide valuable biological insights into past animal life-histories, behaviour and demography, including the effects of environmental change. However, such inferences have often not been possible due to the fragmented nature of the zooarchaeological record and a lack of clear diagnostic skeletal markers. In this study, we test whether the dog (Canis lupus familiaris) nuclear genome is suitable for genetic sex identification in pinnipeds. We initially tested 72 contemporary ringed seal (Pusa hispida) genomes with known sex, using the proportion of X chromosome DNA reads to chromosome 1 DNA reads (i.e. chrX/chr1-ratio) to distinguish males from females. This method was found to be highly reliable, with the ratios clustering in two clearly distinguishable sex groups, allowing 69 of the 72 individuals to be correctly identified according to sex. Secondly, to determine the lower limit of DNA reads required for this method, a subset of the ringed seal genome data was randomly down-sampled. We found a lower threshold of as few as 5000 mapped DNA sequence reads required for reliable sex identification. Finally, applying this standard, sex identification was successfully carried out on a broad set of ancient pinniped samples, including walruses (Odobenus rosmarus), grey seals (Halichoerus grypus) and harp seals (Pagophilus groenlandicus). All three species showed clearly distinct male and female chrX/chr1 ratio groups, providing sex identification of 42–98% of the samples, depending on species and sample quality. The approach described in this study should aid in untangling the putative effects of human activities and environmental change on populations of pinnipeds and other animal species.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".