Who is stressed? Comparing cortisol levels between individuals
Bibliographic record
Abstract
UNLABELLED: Cortisol is the most commonly used biomarker to compare physiological stress between individuals. Its use, however, is frequently inappropriate. Basal cortisol production varies markedly between individuals. Yet, in naturalistic studies that variation is often ignored, potentially leading to important biases. OBJECTIVES: Identify appropriate analytical tools to compare cortisol across individuals and outline simple simulation procedures for determining the number of measurements required to apply those methods. METHODS: We evaluate and compare three alternative methods (raw values, Z-scores, and sample percentiles) to rank individuals according to their cortisol levels. We apply each of these methods to first morning urinary cortisol data collected thrice weekly from 14 cycling Mayan Kaqchiquel women. We also outline a simple simulation to estimate appropriate sample sizes. RESULTS: Cortisol values varied substantially across women (ranges: means: 1.9-2.7; medians: 1.9-2.8; SD: 0.26-0.49) as did their individual distributions. Cortisol values within women were uncorrelated. The accuracy of the rankings obtained using the Z-scores and sample percentiles was similar, and both were superior to those obtained using the cross-sectional cortisol values. Given the interindividual variation observed in our population, 10-15 cortisol measurements per participant provide an acceptable degree of accuracy for across-women comparisons. CONCLUSIONS: The use of single raw cortisol values is inadequate to compare physiological stress levels across individuals. If the distributions of individuals' cortisol values are approximately normal, then the standardized ranking method is most appropriate; otherwise, the sample percentile method is advised. These methods may be applied to compare stress levels across individuals in other populations and species.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".