P614 A strategy to solve complexity of physician diagnosis using a mixture of Patient Reports and Provider Reports: A Case Study of Crohn’s Disease Phenotypes
Bibliographic record
Abstract
Abstract Background The IBD Plexus® Registry, established by the US Crohn’s & Colitis Foundation, combines patient and provider reported survey data collected from doctors’ visits and historical data in electronic medical record (EMR) systems. This multi-resource dataset presents data discrepancies in assessing phenotypes of Crohn’s disease (CD). The aim of this analysis is to describe a strategic process to solve data challenges related to CD phenotype definitions based on the IBD Plexus data. Methods This cross-sectional, multicenter, US based study used data from 11/2016 to 06/2020 from the IBD Plexus® Sparc program to assess demographics, symptoms and treatments among CD patients. Data discrepancies between CD phenotype status reported and as described in historical EMRs were identified. Physicians may use the Montreal classification with/without the Paris modification for phenotype evaluation. To resolve this discrepancy, we explored and evaluated several study phenotype definitions: 1) using phenotypes at visits, the registry suggested method; 2) using phenotype and history of fistula/abscess or stricture in EMR; and 3) using definition 2 without anal stricture. The implementation included two steps: 1) severe conditions (penetrating, stricturing or both) were considered irreversible and defined using the data at any time before 30 days after the registry consent date; 2) the inflammatory condition was positive in the absence of any other reported severe condition during the entire study period. Results The frequency results by phenotypes show small differences across definitions (Figure 1). The discrepancy in frequency by definition1 demonstrated the phenotypes recorded at visits contradicted phenotypes in the EMR. For instance, 0.1%-3.2% or 0.1–0.7% of CD inflammatory patients had subtypes of stricture and subtypes of fistula/abscess, respectively. About 0.5% of CD stricturing patients had intra-abdominal abscess or other fistula (Figure 2). Among CD penetrating patients, 32.0% had history of ileal stricture (Figure 3). Including EMR phenotype variables in definition 2, all discrepancies were resolved. With verification, anal strictures are due to perianal disease which should not be used in the stricturing definition; therefore, the anal stricture was exempt from definition 3. All subtype phenotypes showed 0% discrepancy with study phenotype. A small percentage of positive anal stricture patients was allowed in definition 3. Figure 1. Figure 2. Figure 3. Figure 4. Conclusion When handling the mixture of patient reported and provider-reported data, data discrepancies have many causes. Clarifying the clinical rationale is a key process to resolve discrepancies and accurately define measures of interest.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.061 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".