Comparative structural insights of X protein across species and the “lost” BH3-like domain that may explain the absence of hepatocellular carcinoma in birds
Bibliographic record
Abstract
Chronic hepatitis B virus (HBV) infection is a leading cause of hepatocellular carcinoma (HCC). HBV produces an X protein (HBx) that is heavily linked to HCC, a similar finding seen across the Orthohepadnaviruses. There is an analogous, but truncated form of the X protein in Avihepadnaviruses, but no HCC is seen in infected avian species. This raises questions about the differences between mammalian and avian X proteins and their potential roles in carcinogenesis. Here we explore this question from a structural perspective of X proteins across mammalian and avian hepadnaviruses to determine whether there are any structural features linking these interesting observations. We first compile sequences to create consensus sequences with which to align with each other and subsequently to input into modeling programs, RoseTTAFold (RF) and AlphaFold3 (AF3). Comparative analyses show that mammalian X proteins are longer (154aa and 141aa, respectively for HBx and WHx) and have distinct domains in their C-terminal region, including a BH3-like domain that has been linked with many cancer pathways. Avian X proteins, however, are truncated with a similarly located stop codon. This truncation rids the protein of the BH3-like domain. We sought to find this domain by theoretical read-through or frameshifts and indeed, an analogous BH3-like domain is found that can similarly bind the BH3 target Bcl-2. Based on this work, it may be evolutionarily plausible that a single-nucleotide insertion led to the stop codon and BH3-domain loss that may account for the lack of cancers seen with Avihepadnaviruses.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".