Data and Code: "Historic logbooks reveal spatial footprints of commercial whaling"
Bibliographic record
Abstract
Dataset 1. Transcriptions of the 161 unpublished whaling voyages new to this study.Dataset 2. Bowhead whaling strike data digitized from map figures and tables in Reeves et al. 1983. “Distribution and Migration of the Bowhead Whale, Balaena mysticetus, in the Eastern North American Arctic”. <i>Arctic</i> <b>36</b>:1, 5-64. https://www.jstor.org/stable/40509468Dataset 3. Bowhead whaling strike data digitized from map figures in Ross and McIver 1982. “Distribution of the kills of bowhead whales and other sea mammals by Davis Strait whalers, 1829-1910”. Unpublished manuscript. Arctic Pilot Project, Petro Can., 55D-6th Ave., S.W. Calgary, Alberta, TIP-I44, Canada. 75 pages.Dataset 4. Standalone strike data of bowhead whales and other marine mammals from various other references. We treated strikes labelled “blackfish” as referring to bowhead whales Data from Townsend (1935) and Census of Marine Life were downloaded from the New Bedford Whaling Museum on 8 December 2022.Dataset 5. Cleaned bowhead whale strike data from Datasets 2-4, where we removed data with imprecise spatial and/or temporal descriptions.Dataset 6. Reconstructed whaling voyages in the Bering-Chukchi-Beaufort stock.Dataset 7. Reconstructed whaling voyages in the East Canada-West Greenland stock (including Hudson Bay).Dataset 8. Reconstructed whaling voyages in the East Greenland-Svalbard-Barents stock.Dataset 9. Input data for the Bayesian Structural Causal Model. X- and Y- coordinates are projected using the following CRS: "+proj=stere +lat_0=90 +lat_ts=71 +lon_0=-120 +datum=WGS84 +units=km +no_defs +ellps=WGS84".Software 1. R code for running the Bayesian hDCRW model using Rstan.Software 2. Stan code for the hierarchical Bayesian DCRW model.Software 3. R code for running the Bayesian Structural Causal Model.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.066 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".