Enhancing the utility of International Journal of Epidemiology cohort profiles
Bibliographic record
Abstract
The cohort profiles published by International Journal of Epidemiology (IJE) are intended to support scientific collaboration and enhance the use of longitudinal cohort studies. This welcome addition to other types of IJE articles is very much in line with recent reports1–4 emphasizing the need to maximize substantial research investments already made in large cohort datasets. We examined abstracts of all 45 cohort profiles and nine profile updates published by IJE in 2015 and 2016, and reviewed full articles for 55 cohort profiles along with two profile updates published by IJE in 2017. For abstracts from 2015 and 2016, our primary interest was whether or not authors identified population health intervention studies as an area for collaborative research; 74% (n = 40) of the 54 profile authors indicated they would welcome some sort of collaboration. However, the predominant suggestions for collaboration involved using the cohorts in studies of aetiology, prognosis, risk factors and determinants, and genome-wide associations. Few authors described the use of cohorts for either clinical or population health intervention studies, and none of these profiles described the use of cohort data to evaluate policy interventions. Our review of profile articles for 2017 indicates some increased reporting of intervention studies embedded within cohorts and the use of cohorts to review the impact of policies. The purposes outlined for establishing these more recently reported cohorts included reference to clinical or population health interventions for 35.1 % (n = 20). Seven of these 20 papers indicated that the primary or secondary purposes of the cohort were to inform interventions or the development of assessment tools, rather than to use the cohort as a source of data or platform to examine clinical or population health interventions. Examples of interventions being tested or proposed for examination included vaccines, lead and injury hazard controls, infant formula, antiretroviral therapy, hormone replacement therapy, food fortification policies and national child health policies such as child support and access to education and potable water. Various intervention study designs were described, including randomized controlled trials, historical retrospective policy studies and cross-jurisdictional comparative studies. Additionally, 60% (n = 34) of the 2017 profiles described data linkages already established, and another 32% (n = 18) had linkages planned. Cohorts were linked to a variety of datasets including administrative databases (health, education, occupational and social); disease, vaccine and drug registries; and birth, death, migration and incarceration records. These are encouraging signs that authors are considering the utility of cohorts for intervention studies. We believe this information could be further strengthened with a few additions to the cohort profiles. Therefore, we suggest that journal editors add a descriptive sub-heading(s) to prompt authors to provide more details about the application of their cohort for the purposes of clinical and population health intervention studies. Specifically, authors should be asked to: (i) indicate whether or not their cohort has already been used for clinical or population health intervention research studies and to provide examples; (ii) briefly describe what potential the cohort has for either clinical or population health interventions; and (iii) indicate whether or not linking variables are included in the dataset which might make them particularly amenable to studies of population health interventions such as an examination of distinctive policies in comparable jurisdictions. This would help prompt both authors and potential collaborators to consider the intervention research potential of their cohort datasets. In addition, it would be useful for journals to invite letters and commentaries asking researchers to indicate how selected cohort studies could be optimized for clinical and population health intervention research studies. This work was supported by the Canadian Institutes of Health Research [grant number: 122510]. Conflict of interest: None declared.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.079 | 0.314 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.006 | 0.005 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.002 | 0.005 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".