Comparison of genomes and proteomes of four whole genome-sequenced Campylobacter jejuni from different phylogenetic backgrounds
Bibliographic record
Abstract
Whole genome sequencing (WGS) has been used to assess the phylogenetic relationships, virulence and metabolic differences, and the relationship between gene carriage and host or niche differentiation among populations of C. jejuni isolates. We previously characterized the presence and expression of CJIE4 prophage proteins in four C. jejuni isolates using WGS and comparative proteomics analysis, but the isolates were not assessed further. In this study we compare the closed, finished genome sequences of these isolates to the total proteome. Genomes of the four isolates differ in phage content and location, plasmid content, capsular polysaccharide biosynthesis loci, a type VI secretion system, orientation of the ~92 kb invertible element, and allelic differences. Proteins with 99% sequence identity can be differentiated using isobaric tags for relative and absolute quantification (iTRAQ) comparative proteomic methods. GO enrichment analysis and the type of artefacts produced in comparative proteomic analysis depend on whether proteins are encoded in only one isolate or common to all isolates, whether different isolates have different alleles of the proteins analyzed, whether conserved and variable regions are both present in the protein group analyzed, and on how the analysis is done. Several proteins encoded by genes with very high levels of sequence identity in all four isolates exhibited preferentially higher protein expression in only one of the four isolates, suggesting differential regulation among the isolates. It is possible to analyze comparative protein expression in more distantly related isolates in the context of WGS data, though the results are more complex to interpret than when isolates are clonal or very closely related. Comparative proteomic analysis produced log2 fold expression data suggestive of regulatory differences among isolates, indicating that it may be useful as a hypothesis generation exercise to identify regulated proteins and regulatory pathways for more detailed analysis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".