EXamining ouTcomEs in chroNic Disease in the 45 and Up Study (the EXTEND45 Study): Protocol for an Australian Linked Cohort Study
Bibliographic record
Abstract
BACKGROUND: Chronic kidney disease (CKD) and diabetes are the major causes of death and disability worldwide. They are associated with high health service utilization persisting over many years. Their slow progression and wide clinical variation make them eminently suitable for study in population-based cohorts. However, current understanding of their prevalence, incidence, and progression is largely based on studies conducted in clinical populations. OBJECTIVE: This study aims to establish a novel link between an existing population-based cohort (the 45 and Up Study) and routinely collected laboratory and administrative data to facilitate research across the full disease spectrum of CKD and diabetes. METHODS: In the EXTEND45 Study (EXamining OuTcomEs in chroNic Disease in the 45 and Up Study), baseline questionnaire responses of over 260,000 participants of the 45 and Up Study aged ≥45 years living in New South Wales (NSW), collected between January 2006 and December 2009, are linked to data from laboratory service providers as well as national- and state-based administrative datasets via probabilistic linkage. Routinely collected data were obtained for participants who could be linked between January 2005 and July 2013. Laboratory data will enable the identification of early cases of chronic disease and the assessment of clinically relevant biochemical targets during the disease course. Health administrative datasets will allow for the examination of health service use, pharmacological management, and clinical outcomes. RESULTS: The study received ethics approval from the NSW Population and Health Services Research Ethics Committee in February 2014. Data linkage for 267,153 of the 45 and Up Study participants was completed in June 2016, with congruent linkage achieved for 265,086 (99.23%) individuals. To date, the CKD and diabetes cohorts have been identified (published elsewhere), and a diverse portfolio of research projects relating to disease burden, risk factors, health outcomes, and health service utilization is in development. CONCLUSIONS: The EXTEND45 Study represents an unparalleled opportunity to perform extensive research into diseases of considerable public health and clinical importance. Strengths include the population-based nature of the cohort and the availability of longitudinal information on the complete disease pathway for affected individuals. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): RR1-10.2196/15646.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.047 | 0.033 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.004 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.024 | 0.007 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".