MétaCan
Menu
Back to cohort
Record W4416531244 · doi:10.1111/ppe.70098

Diagnosis Code to Function: Tailoring an Algorithm for Children With Neurodisability

2025· article· en· W4416531244 on OpenAlexaffabout
Katherine Nelson

Bibliographic record

VenuePaediatric and Perinatal Epidemiology · 2025
Typearticle
Languageen
FieldMedicine
TopicCerebral Palsy and Movement Disorders
Canadian institutionsInstitute for Clinical Evaluative SciencesHospital for Sick ChildrenUniversity of Toronto
Fundersnot available
KeywordsIdentification (biology)Field (mathematics)Service (business)TraitHealth careHealth servicesCode (set theory)

Abstract

fetched live from OpenAlex

In 2001, the World Health Organization published the International Classification of Functioning, Disability, and Health (ICF) framework, which expanded the definition of health beyond disease to capture how individuals function in and engage with their environment and community [1]. This broader conceptualisation revolutionised the field of childhood disability, impacting clinical practice, policy, and service delivery. Transformational changes require data to evaluate effectiveness; routinely collected health and education data have been leveraged for this purpose. A few jurisdictions—Denmark, Scotland, Wales, Australia (New South Wales and Western Australia), and Canada (Manitoba)—have population-wide data cross-linkages between the health and education sectors, allowing for the evaluation of educational outcomes for specific clinical cohorts. With the development of the Education and Child Health Insights from Linked Data (ECHILD) database [2], England has joined their ranks. In this issue of Paediatric and Perinatal Epidemiology, Zylbersztejn and colleagues [3] describe the development and thoughtful evaluation of a diagnosis-code-based algorithm to identify children with neurodisability who have functional impairments resulting from neurological conditions. Utilizing linked data from ECHILD, they showed that children with hospital-diagnosed neurodisability have higher healthcare utilisation, mortality, and special educational needs than their peers. Creating new diagnosis-code-based algorithms, such as those for neurodisability, requires balancing sensitivity and specificity to ensure that the algorithm maximises identification of affected children without compromising the likelihood that an identified child has the trait of interest. This commentary articulates the challenges in operationalising terms like ‘neurodisability’ using diagnosis codes, highlighting how key decisions in development impact the final algorithm's performance. In paediatrics, the rarity of many medical conditions limits the feasibility of creating diagnosis-specific cohorts for research. One solution is the creation of ‘non-categorical’ definitions for childhood conditions that group individual diagnoses (such as autism, cerebral palsy, and hearing impairment) into a combined larger population (‘children with neurodisability’) based on a common trait (functional impairment from a neurologic condition). Operationalising a non-categorical definition with diagnosis codes allows researchers to efficiently gather baseline data about the population and to evaluate subsequently developed interventions. For example, ‘children with medical complexity’ (CMC) were defined as those with one or more complex chronic conditions leading to severe functional limitations, increased healthcare utilisation, and substantial support needs [4]. This definition and its operationalisation facilitated the creation and evaluation of complex care clinical programs. ‘Severe neurologic impairment’ (SNI) and ‘neurodisability’ are both non-categorical terms describing children with neurological conditions who have functional limitations. Severe neurologic impairment focuses on children requiring ‘much assistance’ with activities of daily living [5] and has been operationalised in the International Classification of Diseases, 10th revision (ICD-10) [6]. However, SNI's high threshold for impairment was not suitable for this study's goal—predicting population-level special education needs for children with neurodisability—so the authors created a more inclusive algorithm to capture children with any functional limitation arising from a neurologic condition [3]. Function is challenging to operationalise using health administrative data for several reasons. First, function is often not explicitly assessed in clinical encounters. Second, even when it is measured, function is poorly captured by ICD-10 codes, a fact that influenced the development of the ICF 25 years ago. In an ideal world, healthcare data would utilise an ICF-informed coding structure, but in the meantime, operationalising neurodisability requires assuming the degree of functional impacts of individual diagnoses. Third, the ICD-10 includes many diagnoses (such as hypoxic ischemic encephalopathy) with widely variable functional implications. Creating a diagnosis-code-based algorithm requires a binary assessment (include or exclude) for each diagnosis, which subsequently misclassifies children whose functional limitations are greater (for excluded diagnoses) or less (for included diagnoses) than expected. The researchers acknowledge the poor capture of function in administrative data, as well as the ICD-10's biomedical bias [3]. To account for the spectrum of possible functional outcomes, they established a prevalence threshold for inclusion: if more than 50% of children with a diagnosis were anticipated to have a functional impairment, the diagnosis would be included [3]. While this standard is imprecise, given the researchers' goal of identifying a broad cross-section of children who potentially require special educational supports, they reasonably chose to prioritise sensitivity when classifying diagnoses. Frequently, as in this study, the initial pool of diagnosis codes to be assessed for inclusion in a diagnosis-code-based algorithm is drawn from algorithms developed for other purposes. Some studies expand the list of potential codes through targeted review of ICD-10 subsections or by assessing diagnosis codes used in practice among children with the trait of interest. Typically, the goal is comprehensiveness, maximising the number of potentially relevant codes to be assessed. However, the more inclusive the list for review, the more challenging the reviewer's job. Even with clear criteria, adjudicating unfamiliar or highly variable diagnoses is subjective. Requiring consensus from a group of expert reviewers is an important counterbalance; however, there is an unavoidable tension between maximising the likelihood that the algorithm will capture as many affected children as possible and ensuring that the algorithm is appropriately discriminative for the common trait. In future iterations, the neurodisability algorithm could be further expanded by creating a cohort of children who receive the most intensive educational supports, then stratifying the cohort by their neurodisability status using the algorithm. Reviewing diagnosis codes of children classified as ‘no neurodisability’ would allow identification of false negatives—children who carry diagnosis codes associated with neurodisability that are not included in the algorithm—which could then be added to the algorithm. There is an important downside to enriching the algorithm this way: it might exaggerate the relative frequency of children with greater needs when subsequently applied to a general population. Whether such amplification is problematic depends on the appropriate balance of sensitivity and specificity for the study at hand. Another key decision lies in choosing the population for which to apply the algorithm. The assessed health records can be limited to hospitalizations or can also include outpatient encounters; this decision significantly impacts the identified cohort. In this study, outpatient populations would be expected to have a lower prevalence of neurologic conditions and less significant functional impairments (as more significant impairments are associated with a higher risk of hospitalisation). While applying the algorithm to hospitalized patients improves algorithm performance because the positive predictive value increases with a higher baseline prevalence, it likely also amplifies the proportion of children with more intensive special education needs due to greater impairments among hospitalized children. The choice to prioritise specificity (via exclusion of outpatient records) versus sensitivity depends on the preferred direction of bias in the outcome. For this study, with its goal of supporting educational planning, the greater risk would be in underestimating the intensity of population needs; therefore, limiting it to inpatient records is reasonable. The authors appropriately acknowledge the potential limitation of missing some children with likely milder impairments who have not been hospitalized. The goal of this commentary is to make explicit the many nuanced decisions in the development of diagnosis-code-based algorithms that influence which children are ultimately identified by the algorithm. These individually small decisions have a large impact on the final cohort, as a study comparing cohorts ascertained by three different algorithms for paediatric medical complexity nicely demonstrated [7]. In that study, each algorithm identified a different group of children, with limited overlap: 58% of children were identified by only one algorithm, and 12% were identified by all three algorithms. Interestingly, however, the implications of this variability differed depending on the type of outcome. Prevalence, mortality, and health services utilisation outcomes were relatively consistent across the three cohorts, but more specific clinical outcomes (such as the most commonly affected body system) differed substantially. Ultimately, the reliability and usefulness of diagnosis-code-based algorithms depend on the specifics of what they are being used to estimate and the risks of misclassification. Health services outcomes, especially when assessing trends over time in a single population, may be reasonably robust to the inevitable uncertainty inherent in these algorithms. While clinical validation studies are always beneficial for highlighting key performance issues [8], this study's process of external validation—comparing the performance of this neurodisability algorithm to similar algorithms—is likely adequate for its goals. However, future researchers must be cautious in assuming that this or any other algorithm accurately represents a specific clinical cohort. Algorithm creation is a complex process that requires iterative decisions guided by the specific purpose for which the algorithm is being developed. Although the common trait and its definition might be the same, like a bespoke suit requiring alterations for a new owner, the most well-crafted algorithms still require verification and validation when applied in a new context. The author takes full responsibility for this article. The author received no specific funding for this work. The author utilised OpenEvidence to identify potential references for this commentary. The initial draft was written independently by the author. Microsoft Copilot (2025 version) was used in revisions with prompts instructing it to act as a peer reviewer, highlighting strengths and weaknesses, as well as to propose places where the manuscript could be shortened. The author selectively used these suggestions as guidance during manuscript revision and takes full responsibility for the accuracy of the final content. The author declares no conflicts of interest. Data sharing is not applicable to this article as no datasets were generated or analyzed during the current study.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.216
Threshold uncertainty score0.513

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.029
GPT teacher head0.316
Teacher spread0.288 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes2
Has abstractyes

Explore more

Same venuePaediatric and Perinatal EpidemiologySame topicCerebral Palsy and Movement DisordersFrench-language works237,207