A cell-free DNA metagenomic sequencing assay that integrates the host injury response to infection
Bibliographic record
Abstract
High-throughput metagenomic sequencing offers an unbiased approach to identify pathogens in clinical samples. Conventional metagenomic sequencing, however, does not integrate information about the host, which is often critical to distinguish infection from infectious disease, and to assess the severity of disease. Here, we explore the utility of high-throughput sequencing of cell-free DNA (cfDNA) after bisulfite conversion to map the tissue and cell types of origin of host-derived cfDNA, and to profile the bacterial and viral metagenome. We applied this assay to 51 urinary cfDNA isolates collected from a cohort of kidney transplant recipients with and without bacterial and viral infection of the urinary tract. We find that the cell and tissue types of origin of urinary cfDNA can be derived from its genome-wide profile of methylation marks, and strongly depend on infection status. We find evidence of kidney and bladder tissue damage due to viral and bacterial infection, respectively, and of the recruitment of neutrophils to the urinary tract during infection. Through direct comparison to conventional metagenomic sequencing as well as clinical tests of infection, we find this assay accurately captures the bacterial and viral composition of the sample. The assay presented here is straightforward to implement, offers a systems view into bacterial and viral infections of the urinary tract, and can find future use as a tool for the differential diagnosis of infection.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".