Abstract B071: Automated extraction and provision of electronic health record data from children with cancer to National Childhood Cancer Center (NCCR) cancer registries
Bibliographic record
Abstract
Abstract Introduction: The National Childhood Cancer Registry (NCCR) is a central infrastructure that integrates childhood cancer data from central cancer registries and other sources to enhance access to and utilization of pediatric cancer data. The NCCR currently represents >70% of the United States pediatric population. Registries collect and publish cancer incidence and survival data for public health surveillance. Electronic health records (EHRs) are one data source. EHR data are currently obtained by manual collection at individual hospitals that then transfer data to population-based cancer registries. This study aimed to use automated data collection from hospital EHRs to broaden and standardize data reported to registries. Methods: ExtractEHR is an R based software tool that extracts and formats data derived from EHRs from multiple institutions and vendors for research initiatives. ExtractEHR was installed by the hospital technical team to extract the following EHR components: addresses; demographics; inpatient and clinic visits; laboratory, microbiology, pathology, genomic, and radiology test results; medications orders and administrations; procedures; height and weight; and oncology clinician notes. Patients were identified by the central cancer registry and case lists provided to the hospital team. ExtractEHR extracted data for all patients from time of diagnosis through most recent encounter in the EHR. Data elements were extracted as comma-separated values (csv) files. The hospital clinical lead checked data for completeness and worked with the hospital technical and ExtractEHR teams to iteratively adjust ExtractEHR mapping to ensure complete data extraction. Once considered complete, csv files were transferred to the registry team using a secure file transfer system. The initial pilot was completed for patients with 5 diagnosis types over a 2-year period at Children’s Healthcare of Atlanta (CHOA) and transferred to the Georgia Cancer Registry. Results: Data from 306 patients with astrocytoma, Hodgkin Lymphoma, neuroblastoma, Ewing or osteosarcoma, or leukemia diagnosed at CHOA in 2019 or 2020 were extracted using ExtractEHR. There were 3,015,822 data elements extracted across all EHR components, including 832,006 laboratory test results, 574,462 medication administrations, and 34,756 oncology notes. Both structured data (e.g. laboratory results) and unstructured free text data (e.g. notes) were successfully extracted. Data stored in PDFs in the EHR were not extracted due to storage parameters in the EHR data warehouse. All extracted data were successfully transferred to the registry. Conclusions: ExtractEHR can extract and transfer a broad range of clinical data on children treated for cancer to cancer registries. The data granularity obtained using ExtractEHR is impossible to obtain manually. Work is ongoing to expand to all cancers and implement ExtractEHR at 4 additional hospitals. These data will be deidentified by the registry teams and submitted to the NCCR to expand the comprehensiveness of clinical data included in the NCCR. Citation Format: Tamara P. Miller, Johanna Goderre Jones, Elizabeth Hsu, Edward Krause, Yun Gun Jo, Judy Lee, Kevin Ward, Lynne Penberthy, Richard Aplenc. Automated extraction and provision of electronic health record data from children with cancer to National Childhood Cancer Center (NCCR) cancer registries [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Advances in Pediatric Cancer Research; 2024 Sep 5-8; Toronto, Ontario, Canada. Philadelphia (PA): AACR; Cancer Res 2024;84(17 Suppl):Abstract nr B071.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.062 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.008 | 0.007 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.011 | 0.009 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".