Combining cohorts of prospectively collected and linked data to improve health after release from prison
Bibliographic record
Abstract
BackgroundWhile people who experience incarceration have remarkably poor health profiles, undertaking research to inform evidence-based responses is complicated by difficulties of recruiting people in prison; high rates of socioeconomic marginalisation, study attrition; and legislative and financial barriers to linked data research. There are many advantages to pooling data from multiple studies involving people who experience incarceration, including greater statistical power and geographic and participant heterogeneity. However, there are also challenges that need to be addressed. MethodsWe combined four prospective cohort studies of adults released from prisons in four Australian states. Pre-release interviews, and validated screening assessments, were linked to primary care, medicine dispensing, hospital, alcohol and other drug treatment services, ambulatory mental health, ambulance, corrective services, and death records. Data were harmonised by team members reviewing variable definitions and categories within subject domains. ResultsThe combined cohort consists of 4,232 adults, including 1,544 Indigenous people and 905 women, with a median age of 31 years. Data linkage will enable a median of 9.3 years of prospective follow-up after release from incarceration. Differences in data structures, coding systems between and within datasets, changes over time, and grouping of records belonging to the same event were also addressed. ConclusionThis combined multi-site cohort study is an example of a complex, policy-oriented data linkage project. It was developed to underpin evidence-based, culturally appropriate interventions, health policy and service development for people who were incarcerated. This presentation will discuss the processes and pitfalls experienced while building a multi-sectoral, multi-jurisdictional data linkage project.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".