Evaluation of the quality of clinical data collection for a pan-Canadian cohort of children affected by inherited metabolic diseases: lessons learned from the Canadian Inherited Metabolic Diseases Research Network
Bibliographic record
Abstract
Abstract Background The Canadian Inherited Metabolic Diseases Research Network (CIMDRN) is a pan-Canadian practice-based research network of 14 Hereditary Metabolic Disease Treatment Centres and over 50 investigators. CIMDRN aims to develop evidence to improve health outcomes for children with inherited metabolic diseases (IMD). We describe the development of our clinical data collection platform, discuss our data quality management plan, and present the findings to date from our data quality assessment, highlighting key lessons that can serve as a resource for future clinical research initiatives relating to rare diseases. Methods At participating centres, children born from 2006 to 2015 who were diagnosed with one of 31 targeted IMD were eligible to participate in CIMDRN’s clinical research stream. For all participants, we collected a minimum data set that includes information about demographics and diagnosis. For children with five prioritized IMD, we collected longitudinal data including interventions, clinical outcomes, and indicators of disease management. The data quality management plan included: design of user-friendly and intuitive clinical data collection forms; validation measures at point of data entry, designed to minimize data entry errors; regular communications with each CIMDRN site; and routine review of aggregate data. Results As of June 2019, CIMDRN has enrolled 798 participants of whom 764 (96%) have complete minimum data set information. Results from our data quality assessment revealed that potential data quality issues were related to interpretation of definitions of some variables, participants who transferred care across institutions, and the organization of information within the patient charts (e.g., neuropsychological test results). Little information was missing regarding disease ascertainment and diagnosis (e.g., ascertainment method – 0% missing). Discussion Using several data quality management strategies, we have established a comprehensive clinical database that provides information about care and outcomes for Canadian children affected by IMD. We describe quality issues and lessons for consideration in future clinical research initiatives for rare diseases, including accurately accommodating different clinic workflows and balancing comprehensiveness of data collection with available resources. Integrating data collection within clinical care, leveraging electronic medical records, and implementing core outcome sets will be essential for achieving sustainability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.456 | 0.575 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.009 | 0.019 |
| Science and technology studies | 0.009 | 0.006 |
| Scholarly communication | 0.011 | 0.004 |
| Open science | 0.011 | 0.009 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".