The validation of histological criteria from the IAIH‐PG to distinguish AIH from drug‐induced liver injury
Bibliographic record
Abstract
BACKGROUND AND AIMS: To validate the applicability of the new histological criteria for autoimmune hepatitis (AIH) proposed by the International AIH Pathology Group (IAIH-PG) among Chinese patients with AIH and drug-induced liver injury (DILI). METHODS: The gold standard for diagnosis relied on clinical response: discontinuing treatment without relapse supported DILI, while relapse or ongoing immunosuppressive treatment confirmed AIH. This two-centre retrospective cohort study included inpatients with DILI or AIH from January 2002 to March 2023. Cases that underwent liver biopsy were selected according to inclusion and exclusion criteria. The diagnostic performance of the criteria was assessed by an area under the receiver operating characteristic curve (AUROC). RESULTS: Out of 69 patients: AIH (41, 59%) and DILI (28, 41%). The accuracy, sensitivity and specificity of the new histological criteria for likely and possible AIH were 70%, 98% and 29%, respectively, with an AUROC of 0.8236 [95% confidence interval (CI): 0.7533-0.8938]. For likely AIH, the accuracy, sensitivity and specificity were 73%, 61% and 89%, respectively, with an AUROC of 0.9177 [95% CI: 0.8757-0.9596]. Moreover, for possible AIH, significant differences were found in serum alanine aminotransferase levels [178.4 (87.0, 435.0) versus 536.5 (206.9, 930.4) U/L] and antinuclear antibody (ANA) ≥1:160 [10 (67%) versus 1 (6%)], as well as in lobular lymphoplasmacytic infiltrate [15 (100%) versus 12 (71%)] and more than mild inflammation [13 (87%)versus 6 (35%)] between AIH and DILI (all P values were <0.05). CONCLUSION: The new histological criteria exhibit good diagnostic efficacy in distinguishing AIH from DILI in China, with high AUROC. Key discriminators include low aminotransferase, ANA ≥1:160, lobular lymphoplasmacytic infiltrate and more than mild inflammation, which may further improve diagnostic accuracy for AIH.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".