Linking cohort data with administrative health data to develop a new hypertension prediction model to aid precision health approach
Bibliographic record
Abstract
IntroductionHypertension is a common medical condition, affecting 1 in 5 Canadians, and is a major risk factor for heart attack, stroke, and kidney disease. Predicting the risk of developing incident hypertension may help to inform targeted preventive strategies.
 Objectives and ApproachIdentification of major risk factors and incorporation into a multivariable model for risk stratification may help to identify individuals who are at highest risk for developing incident hypertension and would potentially benefit most from intervention. The goal of the proposed research is to develop a robust hypertension prediction model for the general population using the Alberta Tomorrow Project (ATP) cohort data linked with Alberta’s administrative health data. ATP is Alberta's largest population health cohort, contains baseline data on socio-demographic characteristic, personal and family history of disease, medication use, lifestyle and health behavior, environmental exposures, physical measures and bio samples.
 ResultsAlberta’s administrative health data additionally provides information on health care utilization, enrollment, drugs, physician services, and hospital services. A prediction model for hypertension will be developed using logistic regression where information on candidate variables for the model will be gathered from ATP data and outcome (incident hypertension) will be ascertained from administrative health data (physicians/practitioner claim data and hospital discharge abstract data). Lacking follow-up information in current ATP data has laid the foundation of linking the two data sources through an anonymous unique person identifier (e.g. PHN) that will eventually provide follow-up information on ATP participants who are free of hypertension at baseline developed the disease as well as information on other potential variables.
 Conclusion/ImplicationsThe proposed prediction model will help to identify individuals at highest risk for developing hypertension and those who may benefit most from targeted healthy behavioral interventions and/or treatment. Such identification of high risk people may help prevent hypertension as well as the continuing costly cycle of managing hypertension and its complications.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.006 |
| Open science | 0.005 | 0.004 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".