Molecular landscape of kidney allograft tissues data integration portal (NephroDIP): a curated database to improve integration of high-throughput kidney transplant datasets
Bibliographic record
Abstract
Introduction: Kidney transplantation is the optimal treatment for end-stage kidney disease; however, premature allograft loss remains a serious issue. While many high-throughput omics studies have analyzed patient allograft biospecimens, integration of these datasets is challenging, which represents a considerable barrier to advancing our understanding of the mechanisms of allograft loss. Methods: To facilitate integration, we have created a curated database containing all open-access high-throughput datasets from human kidney transplant studies, termed NephroDIP (Nephrology Data Integration Portal). PubMed was searched for high-throughput transcriptomic, proteomic, single nucleotide variant, metabolomic, and epigenomic studies in kidney transplantation, which yielded 9,964 studies. Results: From these, 134 studies with available data detailing 260 comparisons and 83,262 molecules were included in NephroDIP v1.0. To illustrate the capabilities of NephroDIP, we have used the database to identify common gene, protein, and microRNA networks that are disrupted in patients with chronic antibody-mediated rejection, the most important cause of late allograft loss. We have also explored the role of an immunomodulatory protein galectin-1 (LGALS1), along with its interactors and transcriptional regulators, in kidney allograft injury. We highlight the pathways enriched among LGALS1 interactors and transcriptional regulators in kidney fibrosis and during immunosuppression. Discussion: NephroDIP is an open access data portal that facilitates data visualization and will help provide new insights into existing kidney transplant data through integration of distinct studies and modules (https://ophid.utoronto.ca/NephroDIP).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.027 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.023 | 0.016 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.006 | 0.004 |
| Open science | 0.004 | 0.010 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.013 | 0.007 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".