Using Electronic Health Record Data to Rapidly Identify Children with Glomerular Disease for Clinical Research
Bibliographic record
Abstract
Significance Statement Clinical advances in glomerular disease have been stymied by the rarity of these health conditions, making identification of sufficient numbers of patients with glomerular disease for enrollment in research studies challenging, particularly in the pediatric setting. We leveraged the PEDSnet pediatric health system population of >6.5 million children to develop and evaluate a highly sensitive and specific electronic health record (EHR)–based computable phenotype algorithm to identify the largest cohort of children with glomerular disease to date. This tool for rapid cohort identification applied to a robust resource of multi-institutional longitudinal EHR data offers great potential to enhance and accelerate comparative effectiveness and health outcomes research in glomerular disease. Background The rarity of pediatric glomerular disease makes it difficult to identify sufficient numbers of participants for clinical trials. This leaves limited data to guide improvements in care for these patients. Methods The authors developed and tested an electronic health record (EHR) algorithm to identify children with glomerular disease. We used EHR data from 231 patients with glomerular disorders at a single center to develop a computerized algorithm comprising diagnosis, kidney biopsy, and transplant procedure codes. The algorithm was tested using PEDSnet, a national network of eight children’s hospitals with data on >6.5 million children. Patients with three or more nephrologist encounters ( n =55,560) not meeting the computable phenotype definition of glomerular disease were defined as nonglomerular cases. A reviewer blinded to case status used a standardized form to review random samples of cases ( n =800) and nonglomerular cases ( n =798). Results The final algorithm consisted of two or more diagnosis codes from a qualifying list or one diagnosis code and a pretransplant biopsy. Performance characteristics among the population with three or more nephrology encounters were sensitivity, 96% (95% CI, 94% to 97%); specificity, 93% (95% CI, 91% to 94%); positive predictive value (PPV), 89% (95% CI, 86% to 91%); negative predictive value, 97% (95% CI, 96% to 98%); and area under the receiver operating characteristics curve, 94% (95% CI, 93% to 95%). Requiring that the sum of nephrotic syndrome diagnosis codes exceed that of glomerulonephritis codes identified children with nephrotic syndrome or biopsy-based minimal change nephropathy, FSGS, or membranous nephropathy, with 94% sensitivity and 92% PPV. The algorithm identified 6657 children with glomerular disease across PEDSnet, ≥50% of whom were seen within 18 months. Conclusions The authors developed an EHR-based algorithm and demonstrated that it had excellent classification accuracy across PEDSnet. This tool may enable faster identification of cohorts of pediatric patients with glomerular disease for observational or prospective studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.088 | 0.221 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".