The Cortisol Assessment List (CoAL) A tool to systematically document and evaluate cortisol assessment in blood, urine and saliva
Bibliographic record
Abstract
Background: The reliable assessment of cortisol is a necessary requirement to produce replicable research. Several recommendations to increase cortisol assessment reliability exist. However, cortisol assessment methodology is still rather heterogeneous. For this reason, the Cortisol Assessment List (CoAL) was created.The CoAL can be used to guide researchers during the planning phase and document which measures were taken to increase cortisol data reliability in original studies. Moreover, the CoAL can be used to evaluate data quality in meta research. The items representing strategies to obtain reliable cortisol data can be weighted to indicate which are absolutely necessary to consider and which could be applied less restrictively in order to balance data quality and feasibility. In this paper, the construction process of the CoAL is described. Methods: Item synthesis of the CoAL included a literature search to extract empirically based suggestions regarding the reliable assessment of cortisol. Estimates for the item weighting system were obtained by inviting experts in the field to participate in an online survey (n = 25). Inter-rater reliability (IRR) of the CoAL, was determined by letting independent raters use the CoAL to evaluate a set of randomly selected original studies (k = 90). Results: (Cortisol Awakening Response (CAR): 52%; basal cortisol: 52%; reactive cortisol: 44%) in order to obtain reliable cortisol data. Inter-rater agreement was very high (Cohen's Kappa = .98 - 0.99), indicating sufficient psychometric quality of the CoAL. Discussion: The CoAL is the first tool to systematically plan, document and evaluate cortisol assessment. The survey results indicate that the majority of respondents are aware of essential requirements to increase data reliability. However, results were heterogeneous for some items, highlighting the need to start a process of developing a broad scientific consensus regarding reliable cortisol assessment. The implementation of the CoAL could be a first step in this direction. In conclusion, the CoAL reflects empirical evidence and expert knowledge regarding cortisol assessment and can be used as a flexible tool to plan and document empirical studies or evaluate cortisol data quality in meta research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.036 | 0.085 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.018 | 0.009 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.004 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.016 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".