Next-Generation Sequencing and the Crustacean Immune System: The Need for Alternatives in Immune Gene Annotation
Bibliographic record
Abstract
Next-generation sequencing has been a huge benefit to investigators studying non-model species. High-throughput gene expression studies, which were once restricted to animals with extensive genomic resources, can now be applied to any species. Transcriptomic studies using RNA-Seq can discover hundreds of thousands of transcripts from any species of interest. The power and limitation of these techniques is the sheer size of the dataset that is acquired. Parsing these large datasets is becoming easier as more bioinformatic tools are available for biologists without extensive computer programming expertise. Gene annotation and physiological pathway tools such as Gene Ontology and Kyoto Encyclopedia of Genes and Genomes (KEGG) Orthology enable the application of the vast amount of information acquired from model organisms to non-model species. While noble in nature, utilization of these tools can inadvertently misrepresent transcriptomic data from non-model species via annotation omission. Annotation followed by molecular pathway analysis highlights pathways that are disproportionately affected by disease, stress, or the physiological condition being examined. Problems occur when gene annotation procedures only recognizes a subset, often 50% or less, of the genes differently expressed from a non-model organisms. Annotated transcripts normally belong to highly conserved metabolic or regulatory genes that likely have a secondary or tertiary role, if any at all, in immunity. They appear to be disproportionately affected simply because conserved genes are most easily annotated. Evolutionarily induced specialization of physiological pathways is a driving force of adaptive evolution, but it results in genes that have diverged sufficiently to prevent their identification and annotation through conventional gene or protein databases. The purpose of this manuscript is to highlight some of the challenges faced when annotating crustacean immune genes by using an American lobster (Homarus americanus) transcriptome as an example. Immune genes have evolved rapidly over time, facilitating speciation and adaption to highly divergent ecological niches. Complete and proper annotation of immune genes from invertebrates has been challenging. Modulation of the crustacean immune system occurs in a variety of physiological responses including biotic and abiotic stressors, molting and reproduction. A simple method for the identification of a greater number of potential immune genes is proposed, along with a short introductory primer on crustacean immune response. The intended audience is not the advanced bioinformatic user, but those investigating physiological responses who require rudimentary understanding of crustacean immunological principles, but where immune gene regulation is not their primary interest.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".