Code for analysis of polar bear maternal den abundance and distribution in four regions of northern Alaska and Canada within the Southern Beaufort Sea subpopulation boundary (1982-2015)
Bibliographic record
Abstract
We have archived the derived data files and R/JAGS code for our analysis as a U.S. Geological Survey data release (link ). The code is divided into three R scripts: 1) pbdens_landdens_JWM.r contains R code for fitting hierarchical Bayesian models of polar bear maternal den abundance and distribution for the Southern Beaufort Sea (SBS) subpopulation, 1982-2015. This script requires the installation of JAGS (http://mcmc-jags.sourceforge.net/), and several model files in the JAGS programming language (.bug extension) must be present in the same working directory. There are four required model files: allyears_no_timevar.bug, allyears_timetrend_areablock.bug, allyears_timetrend_areablock.bug, allyears_timevar_areadot.bug, allyears_timevar_areavar.bug, and allyears_timevarblock_areablock.bug, which represent the structure of individual models with various levels of complexity (annual variation in the probability of dens occurring on land and the probability of land dens occurring in each of four study regions). 2) polarbearden_SBS_kde_JWM.r will create 95% kernel density estimates of observed polar bear den locations on land for three periods: 1982-2015, 1982-1999, and 2000-2015. This script includes code for basic plots of the kernel density maps, conversion into raster and shapefile formats, and summary statistics for comparing kernel density estimate values among regions of interest within the SBS. 3) pbdens_SBS_RF_RSF_JWM.r will fit a resource selection function for predicting the probability of a location being used as a polar bear maternal den based on environmental characteristics, by contrasting "used" (observed dens) and "available" (random locations within the 95% kernel density estimate boundary where dens were not observed) for polar bear maternal den locations on land in the SBS. This script also includes code for summarizing model fit, evaluating predictor variable importance, and plotting partial dependence plots characterizing the marginal effect of individual predictor variables on the probability of a location being classified as "Used";. All scripts contain additional documentation describing the data objects used in the analysis, their sources, and the analytical tools that we used. The repository also contains an RData object ('pbdens_SBS_JWM.RData') with all data inputs required for each script.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.264 | 0.115 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".