Characterization of full-length hepatitis C virus sequences for subtypes 1e, 1h and 1l, and a novel variant revealed Cameroon as an area in origin for genotype 1
Bibliographic record
Abstract
In this study, we characterized the full-length genome sequences of seven hepatitis C virus (HCV) isolates belonging to genotype 1. These represent the first complete genomes for HCV subtypes 1e, 1h, 1l, plus one novel variant that qualifies for a new but unassigned subtype. The genomes were characterized using 19-22 overlapping fragments. Each was 9400-9439 nt long and contained a single ORF encoding 3019-3020 amino acids. All viruses were isolated in the sera of seven patients residing in, or originating from, Cameroon. Predicted amino acid sequences were inspected and unique patterns of variation were noted. Phylogenetic analysis using full-length sequences provided evidence for nine genotype 1 subtypes, four of which are described for the first time here. Subsequent phylogenetic analysis of 141 partial NS5B sequences further differentiated 13 subtypes (1a-1m) and six additional unclassified lineages within genotype 1. As a result of this study, there are now seven HCV genotype 1 subtypes (1a-1c, 1e, 1g, 1h, 1l) and two unclassified genotype 1 lineages with full-length genomes characterized. Further analysis of 228 genotype 1 sequences from the HCV database with known countries is consistent with an African origin for genotype 1, and with the hypothesis of subsequent dissemination of some subtypes to Asia, Europe and the Americas.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".