Identifying Hemophagocytic Lymphohistiocytosis using Electronic Health Records and Describing the Impact of Treatment on Outcomes (Preprint)
Notice bibliographique
Résumé
UNSTRUCTURED Title: Identifying Hemophagocytic Lymphohistiocytosis using Electronic Health Records and Describing the Impact of Treatment on Outcomes Authors: Suheyla Ocak MD1, Martin Yi MDA2, Agata Wolochacz BSc2, Ida Mehrdadi BSc2, Ahmed Naqvi MD1, Sumit Gupta MD PhD1-2, Lillian Sung MD PhD1-3, Adam Yan MD MBI1-3 1Division of Haematology/Oncology, The Hospital for Sick Children, 555 University Ave, Toronto, ON, Canada, M5G 1X8 2Child Health Evaluative Sciences, The Hospital for Sick Children, 555 University Ave, Toronto, ON, Canada, M5G 1X8 3Information Management Technology, The Hospital for Sick Children, 555 University Ave, Toronto, ON, Canada, M5G 1X8 ADDRESS FOR CORRESPONDENCE: Adam Yan MD, MBI Division of Haematology/Oncology The Hospital for Sick Children, 555 University Avenue, Toronto, Ontario, M5G1X8, Canada Telephone: 416-813-5287 Fax: 416-813-5979 Email: adam.yan@sickkids.ca RUNNING HEAD: HLH in the electronic health record KEY WORDS: HLH, electronic health record, pediatrics WORD COUNT: Abstract 256; Text 2537; Tables 4; Figures 1; Appendices 1 ABSTRACT Objective: To compare different approaches to using the electronic health record (EHR) to build a cohort of Hemophagocytic Lymphohistiocytosis (HLH) patients, and to evaluate characteristics and outcomes of patients meeting the HLH-2004 diagnostic criteria who received HLH-directed therapies to those who did not. Methods: Three approaches to cohort development in the EHR were taken by identifying patients with: (1) an HLH-specific ICD-10 code, (2) an HLH-specific treatment plan, and (3) meeting the HLH-2004 clinical criteria for diagnosis of HLH. Among patients who met the HLH-2004 criteria, we evaluated the characteristics and outcomes of patients who received HLH-directed therapies to those who did not. HLH treatment was defined as either any chemotherapy, or HLH-specific therapy (dexamethasone, methylprednisolone, anakira, ruxolitinib, cyclosporine, etoposide or emapalumab). Results: We identified 388 patients with possible HLH across the three cohorts. An HLH ICD-10 diagnosis (n=220) and meeting five or more clinical criteria (n=245) were much more common than a HLH treatment plan (n=42). Among the patients meeting HLH-2004 clinical criteria, 193 (79%) received HLH-directed therapy. There was no difference in any specific HLH criteria between those who did and did not receive HLH-directed therapy. In-hospital mortality was very high among both groups and was 15.0% among those who received HLH-directed therapy and 13.5% among those who did not receive HLH-directed therapy. Among 1325 patients with an elevated ferritin and fever, only 252 (19%) met >5 clinical criteria. Conclusions: Constructing HLH cohorts from EHR data is challenging, with diagnosis codes, treatment plans, and clinical criteria each capturing distinct but overlapping populations. INTRODUCTION Hemophagocytic lymphohistiocytosis (HLH) is a syndrome represented by excessive inflammation resulting in organ dysfunction.1 Traditionally, HLH has been classified as primary, where there is a documented or presumed genetic etiology,2,3 or secondary, where excessive inflammation can be attributed to a trigger such as infection, malignancy or a rheumatoloical condition in the absence of a genetic etiology.4 Although there are an increasing number of genes identified responsible for primary HLH,5 there are patients with presumed primary HLH without a known mutation. Further, patients with both primary and secondary HLH may require HLH-directed therapy. Consequently, diagnostic criteria have been developed for HLH (Appendix 1), which often guide decision making about treatment initiation. Criteria were initially proposed in 20042,3, and recently updated in 2024.2 In general, diagnosis of HLH can be made based upon molecular or functional cellular findings consistent with HLH in combination with meeting at least 5 clinical criteria (Table 1). While these criteria have been widely used for diagnosis and treatment decision making, there have been questions raised about the specificity of these criteria, particularly as they relate to secondary HLH.4 Most studies of HLH have focused on small cohorts of participants who have received a clinical diagnosis of HLH. To evaluate the utility of HLH criteria, it may be informative to determine the association of each individual HLH criterion against an overall clinical HLH diagnosis, whether HLH-specific treatments received, or HLH outcome such as mortality. Most studies of HLH patients use small highly curated datasets. Few studies have leveraged electronic health record (EHR) data which may represent broader populations and contain a wider range of clinical information. Leveraging EHR data to facilitate clinical research of HLH requires construction of an accurate patient cohort. Given that EHR data is not entered with the goal of facilitating future downstream research, manually entered data such as diagnosis codes can often be incorrect or missing. Computable phenotypes are machine-evaluable definitions for a given condition developed using EHR data. To our knowledge, no attempt has been made to utilize diverse EHR data to develop a machine-evaluable approach to HLH identification.6–8 The primary objective of this study was to compare various approaches to using EHR data to build a cohort of HLH patients- specifically we compared three approaches: (1) HLH-specific diagnosis code usage, (2) HLH-specific treatment plan application and (3) HLH clinical criteria, among all patients at a pediatric hospital. Our hypothesis was that the use of diagnosis codes and treatment plans would under-capture patients, and the use of clinical criteria would over capture patients. Among patients who met at least five HLH clinical criteria, the secondary objective was to compare characteristics and outcomes among those who received HLH-directed therapy vs those who did not. METHODS This observational study was approved by the Research Ethics Board at The Hospital for Sick Children (SickKids). The requirement for informed consent and assent were waived given the retrospective nature of the study. Data Source The data source was the SickKids Enterprise-wide Data in Azure Repository (SEDAR).9 SEDAR is a curated and validated version of the Epic Clarity database organized by clinically relevant units such as patients, encounters, laboratory tests and medication administrations as examples. This project focused on the following SEDAR tables: patient, diagnoses, cancer treatment plans, hospital encounter, non-hospital encounter, medication administration, prescriptions, laboratory tests, pathology results, flowsheets and notes. Operationalizing HLH Criteria We used the HLH-2004 criteria as these would have been the criteria in place for most of the patient cohort (Table 1). The time window to evaluate clinical criteria were centered on an episode, which was usually an inpatient admission spanning admission to discharge. A criteria was considered met if it occurred at any point during the episode. Five of the criteria were based on the laboratory results table (ferritin, cytopenia, hypertriglyceridemia or hypofibrinogenemia, sCD25 and low NK cell activity). Fever was defined as an oral temperature at least 38.3° C once or 38.0°to 38.2° C for at least one hour.10 Two of the criteria required searching of text. Hemophagocytosis was identified by searching for “h(a)emophagocytosis” or “h(a)emophagocytic” in all pathology results. These reports were manually reviewed to identify true hemophagocytosis. Splenomegaly was identified by searching all notes for any of the following terms “splenomegaly”, “big spleen”, “organomegaly”, and “enlarged spleen”. Negative terms were excluded using the following terms within three words previous to the splenomegaly term: “no”, “none”, “absence”, “without” and “negative”. The number of mentions were too high to manually review each note. Thus, a random sample of 20 notes underwent chart review to validate the approach. We found that 19/20 were correct. One of 20 was incorrect. We considered this satisfactory to proceed. Eligibility Criteria We established three cohorts of HLH “diagnosis” based upon encounters that occurred between June 2, 2018 and May 31, 2025. First, patients with a coded diagnosis of HLH were those with an ICD10 code of D76.1 (hemophagocytic lymphohistiocytosis) or D76.2 (hemophagocytic syndrome, infection-associated). Second, cancer treatment plans are electronic care plans that are a component of Epic’s Beacon oncology module. All chemotherapy at our institution must be ordered within a treatment plan, however treatment plan use is restricted to use by oncologists. To that end, an oncologist would therefore order a drug such as emapalumab within a treatment plan, while a rheumatologist using the same drug would not. We manually identified either treatment plan or protocol display names that included “HLH” or “h(a)emophagocytosis”. Use of a treatment plan was defined as having a HLH specific treatment plan applied in Epic. The third approach consisted of identifying the number of HLH 2004 clinical criteria within an encounter. The HLH-specific encounter was the encounter with the maximum number of criteria. If there were more than one encounter with this number of criteria, the first encounter was selected. For establishment of HLH based on clinical criteria, we identified encounters with at least five criteria within that encounter regardless of encounter length. Procedure We were interested in two measures of whether patients received HLH treatment; these might be administered within or outside of an HLH treatment plan. First, we considered any chemotherapy. Second, we considered HLH-directed therapy which we defined as receipt of any dosage of dexamethasone, methylprednisolone, anakira, ruxolitinib, cyclosporine, etoposi
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,048 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,001 |
| Bibliométrie | 0,003 | 0,007 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,002 | 0,001 |
| Science ouverte | 0,000 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,024 | 0,004 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».