Comparative Genomics of<i>Helicobacter pylori</i>: Analysis of the Outer Membrane Protein Families
Bibliographic record
Abstract
The two complete genomic sequences of Helicobacter pylori J99 and 26695 were used to compare the paralogous families (related genes within one genome, likely to have related function) of genes predicted to encode outer membrane proteins which were present in each strain. We identified five paralogous gene families ranging in size from 3 to 33 members; two of these families contained members specific for either H. pylori J99 or H. pylori 26695. Most orthologous protein pairs (equivalent genes between two genomes, same function) shared considerable identity between the two strains. The unusual set of outer membrane proteins and the specialized outer membrane may be a reflection of the adaptation of H. pylori to the unique gastric environment where it is found. One subfamily of proteins, which contains both channel-forming and adhesin molecules, is extremely highly related at the sequence level and has likely arisen due to ancestral gene duplication. In addition, the largest paralogous family contained two essentially identical pairs of genes in both strains. The presence and genomic organization of these two pairs of duplicated genes were analyzed in a panel of independent H. pylori isolates. While one pair was present in every strain examined, one allele of the other pair appeared partially deleted in several isolates.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".