Development of a Composite Health Index in Children with Cystic Fibrosis: A Pipeline for Data Processing, Machine Learning, and Model Implementation using Electronic Health Records
Bibliographic record
Abstract
Cystic Fibrosis (CF) is a heterogeneous multi-faceted genetic condition that primarily affects the lungs and digestive system. For children and young people living with CF, timely management is necessary to prevent the establishment of severe disease. Modern data capture through electronic health records (EHR) have created an opportunity to use machine learning algorithms to classify subgroups of disease to understand health status and prognosis. The overall aim of this thesis was to develop a composite health index in children with CF. \nAn iterative approach to unsupervised cluster analysis was developed to identify homogeneous clusters of children with CF in a pre-existing encounter-based CF database from Toronto Canada. An external validation of the model was carried out in a historical CF dataset from Great Ormond Street Hospital (GOSH) in London UK. The clusters were also re-created and validated using EHR data from GOSH when it first became accessible in 2021. The interpretability and sensitivity of the GOSH EHR model was explored. Lastly, a scoping review was carried out to investigate common barriers to implementation of prognostic machine learning algorithms in paediatric respiratory care. \nA cluster model was identified that detailed four clusters associated with time to future hospitalisation, pulmonary exacerbation, and lung function. The clusters were also associated with different disease related variables such as comorbidities, anthropometrics, microbiology infections, and treatment history. An app was developed to display individualised cluster assignment, which will be a useful way to interpret the cluster model clinically. The review of prognostic machine learning algorithms identified a lack of reproducibility and validations as the major limitation to model reporting that impair clinical translation. \nEHR systems facilitate point-of-care access of individualised data and integrated machine learning models. However, there is a gap in translation to clinical implementation of machine learning models. With appropriate regulatory frameworks the health index developed for children with CF could be implemented in CF care.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".