MétaCan
Menu
Retour à la cohorte
Enregistrement W4416531244 · doi:10.1111/ppe.70098

Diagnosis Code to Function: Tailoring an Algorithm for Children With Neurodisability

2025· article· en· W4416531244 sur OpenAlexaffabout
Katherine Nelson

Notice bibliographique

RevuePaediatric and Perinatal Epidemiology · 2025
Typearticle
Langueen
DomaineMedicine
ThématiqueCerebral Palsy and Movement Disorders
Établissements canadiensInstitute for Clinical Evaluative SciencesHospital for Sick ChildrenUniversity of Toronto
Organismes subventionnairesnon disponible
Mots-clésIdentification (biology)Field (mathematics)Service (business)TraitHealth careHealth servicesCode (set theory)

Résumé

récupéré en direct d'OpenAlex

In 2001, the World Health Organization published the International Classification of Functioning, Disability, and Health (ICF) framework, which expanded the definition of health beyond disease to capture how individuals function in and engage with their environment and community [1]. This broader conceptualisation revolutionised the field of childhood disability, impacting clinical practice, policy, and service delivery. Transformational changes require data to evaluate effectiveness; routinely collected health and education data have been leveraged for this purpose. A few jurisdictions—Denmark, Scotland, Wales, Australia (New South Wales and Western Australia), and Canada (Manitoba)—have population-wide data cross-linkages between the health and education sectors, allowing for the evaluation of educational outcomes for specific clinical cohorts. With the development of the Education and Child Health Insights from Linked Data (ECHILD) database [2], England has joined their ranks. In this issue of Paediatric and Perinatal Epidemiology, Zylbersztejn and colleagues [3] describe the development and thoughtful evaluation of a diagnosis-code-based algorithm to identify children with neurodisability who have functional impairments resulting from neurological conditions. Utilizing linked data from ECHILD, they showed that children with hospital-diagnosed neurodisability have higher healthcare utilisation, mortality, and special educational needs than their peers. Creating new diagnosis-code-based algorithms, such as those for neurodisability, requires balancing sensitivity and specificity to ensure that the algorithm maximises identification of affected children without compromising the likelihood that an identified child has the trait of interest. This commentary articulates the challenges in operationalising terms like ‘neurodisability’ using diagnosis codes, highlighting how key decisions in development impact the final algorithm's performance. In paediatrics, the rarity of many medical conditions limits the feasibility of creating diagnosis-specific cohorts for research. One solution is the creation of ‘non-categorical’ definitions for childhood conditions that group individual diagnoses (such as autism, cerebral palsy, and hearing impairment) into a combined larger population (‘children with neurodisability’) based on a common trait (functional impairment from a neurologic condition). Operationalising a non-categorical definition with diagnosis codes allows researchers to efficiently gather baseline data about the population and to evaluate subsequently developed interventions. For example, ‘children with medical complexity’ (CMC) were defined as those with one or more complex chronic conditions leading to severe functional limitations, increased healthcare utilisation, and substantial support needs [4]. This definition and its operationalisation facilitated the creation and evaluation of complex care clinical programs. ‘Severe neurologic impairment’ (SNI) and ‘neurodisability’ are both non-categorical terms describing children with neurological conditions who have functional limitations. Severe neurologic impairment focuses on children requiring ‘much assistance’ with activities of daily living [5] and has been operationalised in the International Classification of Diseases, 10th revision (ICD-10) [6]. However, SNI's high threshold for impairment was not suitable for this study's goal—predicting population-level special education needs for children with neurodisability—so the authors created a more inclusive algorithm to capture children with any functional limitation arising from a neurologic condition [3]. Function is challenging to operationalise using health administrative data for several reasons. First, function is often not explicitly assessed in clinical encounters. Second, even when it is measured, function is poorly captured by ICD-10 codes, a fact that influenced the development of the ICF 25 years ago. In an ideal world, healthcare data would utilise an ICF-informed coding structure, but in the meantime, operationalising neurodisability requires assuming the degree of functional impacts of individual diagnoses. Third, the ICD-10 includes many diagnoses (such as hypoxic ischemic encephalopathy) with widely variable functional implications. Creating a diagnosis-code-based algorithm requires a binary assessment (include or exclude) for each diagnosis, which subsequently misclassifies children whose functional limitations are greater (for excluded diagnoses) or less (for included diagnoses) than expected. The researchers acknowledge the poor capture of function in administrative data, as well as the ICD-10's biomedical bias [3]. To account for the spectrum of possible functional outcomes, they established a prevalence threshold for inclusion: if more than 50% of children with a diagnosis were anticipated to have a functional impairment, the diagnosis would be included [3]. While this standard is imprecise, given the researchers' goal of identifying a broad cross-section of children who potentially require special educational supports, they reasonably chose to prioritise sensitivity when classifying diagnoses. Frequently, as in this study, the initial pool of diagnosis codes to be assessed for inclusion in a diagnosis-code-based algorithm is drawn from algorithms developed for other purposes. Some studies expand the list of potential codes through targeted review of ICD-10 subsections or by assessing diagnosis codes used in practice among children with the trait of interest. Typically, the goal is comprehensiveness, maximising the number of potentially relevant codes to be assessed. However, the more inclusive the list for review, the more challenging the reviewer's job. Even with clear criteria, adjudicating unfamiliar or highly variable diagnoses is subjective. Requiring consensus from a group of expert reviewers is an important counterbalance; however, there is an unavoidable tension between maximising the likelihood that the algorithm will capture as many affected children as possible and ensuring that the algorithm is appropriately discriminative for the common trait. In future iterations, the neurodisability algorithm could be further expanded by creating a cohort of children who receive the most intensive educational supports, then stratifying the cohort by their neurodisability status using the algorithm. Reviewing diagnosis codes of children classified as ‘no neurodisability’ would allow identification of false negatives—children who carry diagnosis codes associated with neurodisability that are not included in the algorithm—which could then be added to the algorithm. There is an important downside to enriching the algorithm this way: it might exaggerate the relative frequency of children with greater needs when subsequently applied to a general population. Whether such amplification is problematic depends on the appropriate balance of sensitivity and specificity for the study at hand. Another key decision lies in choosing the population for which to apply the algorithm. The assessed health records can be limited to hospitalizations or can also include outpatient encounters; this decision significantly impacts the identified cohort. In this study, outpatient populations would be expected to have a lower prevalence of neurologic conditions and less significant functional impairments (as more significant impairments are associated with a higher risk of hospitalisation). While applying the algorithm to hospitalized patients improves algorithm performance because the positive predictive value increases with a higher baseline prevalence, it likely also amplifies the proportion of children with more intensive special education needs due to greater impairments among hospitalized children. The choice to prioritise specificity (via exclusion of outpatient records) versus sensitivity depends on the preferred direction of bias in the outcome. For this study, with its goal of supporting educational planning, the greater risk would be in underestimating the intensity of population needs; therefore, limiting it to inpatient records is reasonable. The authors appropriately acknowledge the potential limitation of missing some children with likely milder impairments who have not been hospitalized. The goal of this commentary is to make explicit the many nuanced decisions in the development of diagnosis-code-based algorithms that influence which children are ultimately identified by the algorithm. These individually small decisions have a large impact on the final cohort, as a study comparing cohorts ascertained by three different algorithms for paediatric medical complexity nicely demonstrated [7]. In that study, each algorithm identified a different group of children, with limited overlap: 58% of children were identified by only one algorithm, and 12% were identified by all three algorithms. Interestingly, however, the implications of this variability differed depending on the type of outcome. Prevalence, mortality, and health services utilisation outcomes were relatively consistent across the three cohorts, but more specific clinical outcomes (such as the most commonly affected body system) differed substantially. Ultimately, the reliability and usefulness of diagnosis-code-based algorithms depend on the specifics of what they are being used to estimate and the risks of misclassification. Health services outcomes, especially when assessing trends over time in a single population, may be reasonably robust to the inevitable uncertainty inherent in these algorithms. While clinical validation studies are always beneficial for highlighting key performance issues [8], this study's process of external validation—comparing the performance of this neurodisability algorithm to similar algorithms—is likely adequate for its goals. However, future researchers must be cautious in assuming that this or any other algorithm accurately represents a specific clinical cohort. Algorithm creation is a complex process that requires iterative decisions guided by the specific purpose for which the algorithm is being developed. Although the common trait and its definition might be the same, like a bespoke suit requiring alterations for a new owner, the most well-crafted algorithms still require verification and validation when applied in a new context. The author takes full responsibility for this article. The author received no specific funding for this work. The author utilised OpenEvidence to identify potential references for this commentary. The initial draft was written independently by the author. Microsoft Copilot (2025 version) was used in revisions with prompts instructing it to act as a peer reviewer, highlighting strengths and weaknesses, as well as to propose places where the manuscript could be shortened. The author selectively used these suggestions as guidance during manuscript revision and takes full responsibility for the accuracy of the final content. The author declares no conflicts of interest. Data sharing is not applicable to this article as no datasets were generated or analyzed during the current study.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,001
score de la tête « metaresearch » (Gemma)0,001
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: Observationnel
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,216
Score d'incertitude au seuil0,513

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0010,001
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,029
Tête enseignante GPT0,316
Écart entre enseignants0,288 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeObservationnel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2025
Routes d'admission2
Résumé présentoui

Explorer davantage

Même revuePaediatric and Perinatal EpidemiologyMême sujetCerebral Palsy and Movement DisordersTravaux en français237 207