Factors Influencing the Sensitivity and Specificity of Conventional Sequencing in Human Immunodeficiency Virus Type 1 Tropism Testing
Bibliographic record
Abstract
Human immunodeficiency virus type 1 (HIV-1) V3 loop sequence can be used to infer viral coreceptor use. The effect of input copy number on population-based sequencing of the V3 loop of HIV-1 was examined through replicate deep and population-based sequencing of samples with known tropism, a heterogeneous clinical sample (624 population-based sequences and 47 deep-sequencing replicates), and a large cohort of clinical samples from phase III clinical trials of maraviroc including the MOTIVATE/A4001029 studies (n = 1,521). Proviral DNA from two independent samples from each of 101 patients from the MOTIVATE/A4001029 studies was also analyzed. Cumulative technical error occurred at a rate of 3 × 10(-4) mismatches/bp, without observed effect on inferred tropism. Increasing PCR replication increased minority species detection with an ~10% minority population detected in 18% of cases using a single replicate at a viral load of 1,072 copies/ml and in 44% of cases using three replicates. The nucleotide prevalence detected by population-based and deep sequencing were highly correlated (Spearman's ρ, 0.73), and the accuracy increased with increasing input copy number (P < 0.001). Triplicate sequencing was able to predict tropism changes in the MOTIVATE/A4001029 studies for both low (P = 0.05) and high (P = 0.02) viral loads. Sequences derived from independently extracted and processed samples of proviral DNA for the same patient were equivalent to replicates from the same extraction (P = 0.45) and had correlated position-specific scoring matrix scores (Spearman's ρ, 0.75; P << 0.001); however, concordance in tropism inference was only 83%. Input copy number and PCR replication are important factors in minority species detection in samples with significant heterogeneity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.035 | 0.074 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".