Reproducibility Studies on Arteriolar Hyaline Thickening Scoring in Calcineurin Inhibitor-Treated Renal Allograft Recipients
Bibliographic record
Abstract
Arteriolar hyaline thickening (AH) is the most characteristic lesion of chronic calcineurin inhibitor nephrotoxicity. This study was performed to compare the inter-observer reproducibility of AH scoring using Banff criteria and a newly proposed criterion. Forty-five nonprotocol post-transplant biopsies from 38 patients immunosuppressed with tacrolimus or cyclosporine A (CsA) were included. The severity of AH was blindly scored by three observers. According to the new criteria, AH is graded based on circular vs. noncircular involvement and the number of arterioles involved. The kappa statistics were used to assess the inter-observer reproducibility. Twenty-seven (60%) biopsies showed AH. The AH grades by both criteria were correlated with serum creatinine at biopsy and inversely correlated with estimated glomerular filtration rate (GFR) (p < 0.05). The recent AH criteria improved the mean pairwise agreement (79.4% vs. 68%) and the overall kappa value (0.67 vs. 0.52) (p = 0.02) compared to Banff criteria. The mean inter-slide variation using Banff and the new criterion were 23% and 27.6%, respectively (p > 0.05). The new AH criterion results in better inter-observer reproducibility, and is clinically validated against serum creatinine and estimated GFR. There is substantial intra-biopsy variation, therefore, evaluation of more than one section is crucial to determine severity of arteriolar damage more accurately.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.070 | 0.144 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".