MétaCan
Menu
Back to cohort
Record W6930144128 · doi:10.5281/zenodo.10895367

PREreview of "Giant genes are rare but implicated in cell wall degradation by predatory bacteria"

2024· peer-review· en· W6930144128 on OpenAlexaboutno aff

Bibliographic record

VenueZenodo (CERN European Organization for Nuclear Research) · 2024
Typepeer-review
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicBacterial Genetics and Biotechnology
Canadian institutionsnot available
Fundersnot available
KeywordsGiant VirusGeneBacteriaPhylogenetic treeArchaeaOpen reading frameCell wallPeptide sequence

Abstract

fetched live from OpenAlex

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/10895367. We, the students of MICI5029/5049, a Graduate Level Molecular Pathogenesis Journal Club at Dalhousie University in Halifax, NS, Canada, hereby submit a review of the following BioRxiv preprint: Giant genes are rare but implicated in cell wall degradation by predatory bacteria Jacob West-Roberts, Luis Valentin-Alvarado, Susan Mullen, Rohan Sachdeva, Justin Smith, Laura A. Hug, Daniel S. Gregoire, Wentso Liu, Tzu-Yu Lin, Gabriel Husain, Yuki Amano, Lynn Ly, Jillian F. Banfield bioRxiv 2023.11.21.568195; doi: https://doi.org/10.1101/2023.11.21.568195 We will adhere to the Universal Principled (UP) Review guidelines proposed in: Universal Principled Review: A Community-Driven Method to Improve Peer Review. Krummel M, Blish C, Kuhns M, Cadwell K, Oberst A, Goldrath A, Ansel KM, Chi H, O'Connell R, Wherry EJ, Pepper M; Future Immunology Consortium. Cell. 2019 Dec 12;179(7):1441-1445. doi: 10.1016/j.cell.2019.11.029 SUMMARY: Ultra-small bacteria in the phylum Omnitrophota contain some open reading frames (ORFs) greater than 20kbp that should theoretically encode giant proteins (>7000 amino acids). There has been speculation that some features of these giant proteins might allow Omnitrophota to engage in parasitic or predatory activities, but in the absence of isolation and culture of these bacteria, little concrete evidence supports this idea. Here, the authors provide a survey of predicted giant proteins in Omnitrophota and compare the phylogenetic distribution of giant proteins in bacteria and archaea. They found archaea to have fewer giant genes compared to bacteria, and amongst the bacteria, Omnitrophota was found to have the highest average copy number of ORFs predicted to encode protein products greater than 10K amino acids. Most giant protein sequences contained multiple transmembrane domains as well as peptidase domains that suggest localization to the cell membrane and proteolytic activity. Additionally, similarity between Omnitrophota giant genes and those of other predatory bacteria indicate that these proteins might be involved in the degradation of cell walls of prey species. Finally, in silico structural analysis predicted that the encoded giant proteins form novel tubule-like structures of unknown function. This study highlights the need for further study to understand what, if any, utility is provided in possessing giant genes and proteins. OVERALL ASSESSMENT: Reader engagement with this fascinating research topic was undermined by poor connections between the figures and the text, with some figures not discussed in the text at all. We noted several instances where figure annotation and legends could be improved for clarity. Take-home messages should be expressed clearly throughout the manuscript. Our readers would have benefitted from a longer introduction that could provide more context and field-specific evidence for giant proteins and their functions. Considering the large size of the putative ORF, we noted with concern that the authors identified metagenomic assembly quality as a hurdle; accordingly, we wanted to see stronger evidence that the authors had achieved the metagenomic assembly quality necessary to confidently assign the reported giant ORFs. DETAILED U.P. ASSESSMENT: OBJECTIVE CRITERIA (QUALITY) 1. Quality: Experiments (1–3 scale; note: 1 is best on this scale) SCORE = 2.5 · Figure by figure, do experiments, as performed, have the proper controls? Do analyses use the best-possible (most unambiguous) available methods quantified via appropriate statistical comparisons? · In general, the authors provide transparency regarding the methodology and limitations. This aspect of the manuscript could be improved by providing some clarity about the source of the 46 giant genes from the 1873 genomes. The authors described how they curated 3 whole genomes. But of the 46 giant genes, how many came from the curated whole genomes and how many from the metagenomic data? · We have concerns about the methodology used to predict protein structure. For these predictions, the authors split the amino acid sequences into chunks of 1000 residues, at 500 amino acid intervals. This means that there is an overlap of 2 segments to get a sense of the entire protein structure. We would have appreciated an effort to bootstrap this process to improve the "depth" of coverage – maybe 1000 residues at every 100 AA, or a combination of long and short sequences overlapping. In general, we struggled to understand how the authors stitched the sequences together. 2. Quality: Completeness (1–3 scale) SCORE = 2.5 · Does the collection of experiments and associated analysis of data support the proposed title- and abstract-level conclusions? Typically, the major (title- or abstract-level) conclusions are expected to be supported by at least two experimental systems. · The current dataset does not sufficiently support the authors' title- and abstract-level conclusions. While the authors address that metagenomic assembly quality is a hurdle to accurately identifying a giant ORF they do not provide the reader with confidence that they have achieved the metagenomic assembly quality necessary to predict genes of this size. For example, hybrid sequencing was done on a minority of samples and language describing sequence polishing suggests lack of expertise with the subject. · The links between the giant genes, the omnitrophota phylum, and cell wall degradation to support predatory lifestyle, remain underdeveloped. Including hypotheses and model diagrams in the Discussion about how giant genes are thought contribute to bacterial fitness and predation in its ecological niche (given that there has been other papers published about this phylum and the authors collected DNA from environmental samples) may help the reader. See the following paper: https://ami-journals.onlinelibrary.wiley.com/doi/full/10.1111/1462-2920.16170 · Are there experiments or analyses that have not been performed but if ''true'' would disprove the conclusion (sometimes considered a fatal flaw in the study)? In some cases, a reviewer may propose an alternative conclusion and abstract that is clearly defensible with the experiments as presented, and one solution to ''completeness'' here should always be to temper an abstract or remove a conclusion and to discuss this alternative in the discussion section. · Higher quality assembly of this region could reveal stop codons, cleavage sites, or other genetic determinants suggesting this region is not transcribed and translated as a single protein. 3. Quality: Reproducibility (1–3 scale) SCORE = 2 · Figure by figure, were experiments repeated per a standard of 3 repeats or 5 mice per cohort, etc.? · Is there sufficient raw data presented to assess the rigor of the analysis? · More long- and short-read sequencing at higher depths could provide the confidence needed to support the authors' claims. · Newly sequenced genomes are not being uploaded to NCBI until publication, which makes it difficult to assess their analysis. · Are methods for experimentation and analysis adequately outlined to permit reproducibility? · Predicted folded structures calculated using Alphafold2 (see Methods)." There is some description of how they used AlphaFold but they don't say they used it in the methods or cite it. · We appreciate that the authors have made their code available on FigShare. · If a ''discovery' dataset is used, has a ''validation' cohort been assessed and/or has the issue of false discovery been addressed? N/A 4. Quality: Scholarship (1–4 scale but generally not the basis for acceptance or rejection) SCORE = 1.5 · Has the author cited and discussed the merits of the relevant data that would argue against their conclusion? · Has the author cited and/or discussed the important works that are consistent with their conclusion and that a reader should be especially familiar when considering the work? · Specific (helpful) comments on grammar, diction, paper structure, or data presentation (e.g., change a graph style or color scheme) go in this section, but scores in this area should not be significant basis for decisions. · The Introduction is quite brief and does not properly equip the reader to appreciate the major findings of the manuscript or place them in the context of the field, particularly with regard to these bacteria and fitness in their ecological niche (i.e. predation). · We recommend that the authors cite relevant literature throughout the Results and Discussion sections to properly support statements and provide reader with the tools to investigate further. · Headings should be more informative throughout. · Figure-by-figure: · Fig 1: To improve comprehension, we suggest the authors add labels to Fig 1A (the rings are not labelled, so the reader must find this information in the figure legend). Additionally, increasing the font size on the x-axis and providing axis labels for Figures 1B & 1C would make them easier to understand. The scales differ between B & C, which gives the false impression that these values are similar (unless the reader can zoom in to read the x-axis scales). · Figs 2 and 3: That these figures should deleted because they are not mentioned in the Results text. · Fig 7: This figure is mis-referenced in the "Conserved functionalities encoded in proximity to giant genes" section of the Results. MORE SUBJECTIVE CRITERIA (IMPACT): 1. Impact: Novelty/Fundamental and Broad Interest (1–4 scale) SCORE= 2.5 A score here should be accompanied by a statement delineating the most interesting and/or important conceptual finding(s), as they stand right now with the current scope of the paper. A ''1'' would be expected to be understood for the importance by a layperson but would also be of top interest (have lasti

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.023
metaresearch head score (Gemma)0.061
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Evaluation · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Other · Consensus signal: none
Teacher disagreement score0.977
Threshold uncertainty score0.000

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0230.061
Meta-epidemiology (narrow)0.0010.002
Meta-epidemiology (broad)0.0030.003
Bibliometrics0.0070.003
Science and technology studies0.0040.003
Scholarly communication0.0090.005
Open science0.0050.006
Research integrity0.0050.007
Insufficient payload (model declined to judge)0.1070.084

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.025
GPT teacher head0.250
Teacher spread0.226 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainEvaluation
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueZenodo (CERN European Organization for Nuclear Research)Same topicBacterial Genetics and BiotechnologyFrench-language works237,207