Practice makes perfect? Inter‐analyst variation in the identification of fish remains from archaeological sites
Bibliographic record
Abstract
Abstract Identification of faunal specimens based on a morphological comparison with known‐identity reference specimens is the standard methodology used in zooarchaeological analysis. However, the accuracy of identifications is rarely considered. In this paper, we report results of an experiment in which 13 analysts were asked to identify 50 fish skeletal elements from a reference collection and 50 fish skeletal elements from an archaeological collection in southern Ontario. The type and level of experience of the analysts and the amount of time they invested in the identification were controlled. The archaeological specimens were subsequently identified taxonomically using ZooMS. Our findings demonstrate that taxonomic and element identifications are far from perfect, both in the reference collection set and in the archaeological collection set. Probable contributing factors include the richness of taxonomic groups; distinctiveness of skeletal morphology; experience level of the analyst; and size of the individual specimens and whether the analyst had access to comprehensive, well‐labeled reference collections. We recommend emphasis be placed in training on the importance, for most species, of not making a taxonomic identification unless the element identification is certain; conservatism in identification of species in groups with many members; clear knowledge of the range of species possible within a region; and active involvement by the instructor or mentor to ensure that neophyte analysts are corrected.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".