Covariate data for "The association between alcohol consumption per capita and suicide mortality across 30 European countries"
Bibliographic record
Abstract
Contains covariate data for "Association between alcohol consumption per capita and suicide mortality across 30 European countries" which were extracted from the Pew Research Center (pewresearch.org), World Bank Group (worldbank.org), and Eurostat (ec.europa.eu/eurostat). Also contains dummy variables to represent: the 2008 global economic recession, changes from ICD-9 to ICD-10, and the COVID-19 pandemic. All covariates which were initially considered are included in this dataset. However, data were further cleaned according to methods described in the associated publication prior to analysis. Within the dataset: edu = Educational attainment (completion of post-secondary or equivalent); lit = Literacy, adult total (% of people ages 15 and above); unemp = Unemployment, total (% of total labor force) (modeled ILO estimate); divorce = Divorce rate; migration = Net migration rate; relig.muslim = Proportion of the population who identified as Muslim; relig.buddhist = Proportion of the population who identified as Buddhist; lff = Female labour force participation (% of total labor force); gdp = Gross domestic product based on purchasing power parity (GDP (PPP)) gini = Gini index; density = Population density; urban = Proportion of the population living in urban areas; propold = Proportion of the population aged 65+ years of age; recession, covid, icd: Dummy variables detailed above. The relevant citations and attributions are as follows: Liu J. Table: Muslim Population by Country. Published online January 27, 2011. Accessed June 27, 2025. https://www.pewresearch.org/religion/2011/01/27/table-muslim-population-by-country/ Zanetti CH Marcin Stonawski, Yunping Tong, Stephanie Kramer, Anne Shi and Nick. Religious Composition by Country, 2010-2020. Published online June 9, 2025. Accessed July 28, 2025. https://www.pewresearch.org/religion/feature/religious-composition-by-country-2010-2020/ World Bank Group. World Bank Open Data. World Bank Open Data. Accessed June 27, 2025. https://data.worldbank.org World Bank Group. World Development Indicators. DataBank. Accessed June 27, 2025. https://databank.worldbank.org/source/world-development-indicators Eurostat. Divorce indicators. Published online 2022. doi:10.2908/DEMO_NDIVIND Eurostat. Population change - Demographic balance and crude rates at national level. Published online 2022. doi:10.2908/DEMO_GIND
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.005 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.054 | 0.022 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".