Assessing individual and population variability in degenerative joint disease prevalence using generalized linear mixed models
Bibliographic record
Abstract
OBJECTIVES: In this paper, we introduce the use of generalized linear mixed models (GLMM) as a better alternative to traditional statistical methods for studying factors associated to the prevalence of degenerative joint disease (DJD) in bioarchaeological contexts. MATERIALS AND METHODS: DJD prevalence was assessed for the appendicular joints and the spine of a Spanish population dated from the 15th to the 18th century. Data were analyzed using contingency tables, logistic regression models, and logistic GLMM. RESULTS: In general, results from GLMMs find agreement in other methods. However, by being able to analyze the data at the level of individual bones instead of aggregated joints or limbs, GLMMs are capable of revealing associations that are not evident in other frameworks. DISCUSSION: Currently widely available in statistical analysis software, GLMMs can accommodate a wide array of data distributions, account for hierarchical correlations, and return estimates of DJD prevalence within individuals and skeletal locations that are unbiased by the effect of covariates. This gives clear advantages for the analysis of bioarchaeological datasets which can lead to more robust and comparable analyses across populations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.037 | 0.071 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.004 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".