Interchromosomal Segmental Duplications Explain the Unusual Structure of PRSS3, the Gene for an Inhibitor-Resistant Trypsinogen
Bibliographic record
Abstract
Homo sapiens possess several trypsinogen or trypsinogen-like genes of which three (PRSS1, PRSS2, and PRSS3) produce functional trypsins in the digestive tract. PRSS1 and PRSS2 are located on chromosome 7q35, while PRSS3 is found on chromosome 9p13. Here, we report a variation of the theme of new gene creation by duplication: the PRSS3 gene was formed by segmental duplications originating from chromosomes 7q35 and 11q24. As a result, PRSS3 transcripts display two variants of exon 1. The PRSS3 transcript whose gene organization most resembles PRSS1 and PRSS2 encodes a functional protein originally named mesotrypsinogen. The other variant is a fusion transcript, called trypsinogen IV. We show that the first exon of trypsinogen IV is derived from the noncoding first exon of LOC120224, a chromosome 11 gene. LOC120224 codes for a widely conserved transmembrane protein of unknown function. Comparative analyses suggest that these interchromosomal duplications occurred after the divergence of Old World monkeys and hominids. PRSS3 transcripts consist of a mixed population of mRNAs, some expressed in the pancreas and encoding an apparently functional trypsinogen and others of unknown function expressed in brain and a variety of other tissues. Analysis of the selection pressures acting on the trypsinogen gene family shows that, while the apparently functional genes are under mild to strong purifying selection overall, a few residues appear under positive selection. These residues could be involved in interactions with inhibitors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".