Outer membrane pore protein prediction in mycobacteria using genomic comparison
Bibliographic record
Abstract
Proteins responsible for outer membrane transport across the unique membrane structure of Mycobacterium spp. are attractive drug targets in the treatment of human diseases caused by the mycobacterial pathogens, Mycobacterium tuberculosis, M. bovis, M. leprae and M. ulcerans. In contrast with Escherichia coli, relatively few outer-membrane proteins (OMPs) have been identified in Mycobacterium spp., largely due to the difficulties in isolating mycobacterial membrane proteins and our incomplete understanding of secretion mechanisms and cell wall structure in these organisms. To further expand our knowledge of these elusive proteins in mycobacteria, we have improved upon our previous method of OMP prediction in mycobacteria by taking advantage of genomic data from seven mycobacteria species. Our improved algorithm suggests 4333 sequences as putative OMPs in seven species with varying degrees of confidence. The most virulent pathogenic mycobacterial species are slightly enriched in these selected sequences. We present examples of predicted OMPs involved in horizontal transfer and paralogy expansion. Analysis of local secondary structure content allowed identification of small domains predicted to perform as OMPs; some examples show their involvement in events of tandem duplication and domain rearrangements. We discuss the taxonomic distribution of these discovered families and architectures, often specific to mycobacteria or the wider taxonomic class of Actinobacteria. Our results suggest that OMP functionality in mycobacteria is richer than expected and provide a resource to guide future research of these understudied proteins.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".