Assessment of reliability and validity of IBD phenotyping within the National Institutes of Diabetes and Digestive and Kidney Diseases (NIDDK) IBD Genetics Consortium (IBDGC)
Bibliographic record
Abstract
BACKGROUND: The NIDDK IBD Genetics Consortium (IBDGC) collects DNA and phenotypic data from inflammatory bowel disease (IBD) subjects to provide a resource for genetic studies. No previous studies have been performed on the reliability and validity of phenotypic determinations in either Crohn's disease (CD) or ulcerative colitis (UC) using primary records. Our aim was to determine the reliability and validity of these phenotypic assessments. METHODS: The de-identified records of 30 IBD patients were reviewed by 2 phenotypers per center using a standard protocol for phenotypic assessment. Each phenotyper evaluated 10 charts on 2 occasions 5 months apart. Reliability was expressed as the kappa (kappa) statistic. Performance characteristics were determined by comparison to a consensus-derived "gold standard" and by generation of receiver operating characteristic (ROC) curves. RESULTS: Agreement for diagnosis was excellent (kappa = 0.82; 95% confidence interval [CI]: 0.71-0.92). Agreement for CD location was good for jejunal, ileal, colorectal, and perianal disease with kappa between 0.60 and 0.74 but was fair for esophagogastroduodenal (kappa = 0.36). Agreement for UC extent (kappa = 0.67; 95% CI: 0.48-0.85), and CD behavior (kappa = 0.67; 95% CI: 0.49-0.83) were very good. Area under the ROC curves was greater than 0.84 for diagnosis, CD behavior, UC extent, and ileal and colonic CD location. CONCLUSIONS: IBD phenotype classification using a standard protocol exhibited very good to excellent inter- and intrarater agreement and validity. This study highlights the importance of standard protocols in generating reliable and valid phenotypic assessments. The data will facilitate estimates of phenotyping misclassification rates that should be considered when making inferences from IBD genotype-phenotype studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".