Radiocarbon dating of Mesolithic human remains
Bibliographic record
Abstract
This is to introduce what I hope will be an ongoing set of papers in Mesolithic Miscellany. The first of these follows, written in collaboration with Canadian colleagues. I feel that some explanation is required for what is a project in data management. Its starting place is my experience in putting together a recent paper on burial chronology for the proceedings of the MESO2005 conference (Meiklejohn et al. 2009). In that paper we encountered the problem of the highly erratic publication of radiocarbon dates. While clearly reflecting the proprietary nature of these dates, it can make it extremely difficult to obtain a full picture of the state of dating for any given type of archaeological information. To provide context, in the aftermath of presenting a preliminary paper at Belfast, it took over six months of intensive work to gather the information appearing in the paper’s appendix, without which the paper could not have been written. And this length of time was still required even though the database used in the paper had been started over a decade earlier, rooted in a paper that appeared thirty years ago (Newell et al. 1979). The core issue of data availability has been exacerbated over time. The sheer number of dates has accelerated over the years, making any database for a given period or phenomenon of enormous size. One result of this acceleration is that, in a counterproductive way, regular publication of date lists in Radiocarbon is, for practical purposes, finished. Laboratories that ceased publication earliest were generally those producing the
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.006 | 0.007 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.008 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".