Nucleomorph Genome Sequence of the Cryptophyte Alga Chroomonas mesostigmatica CCMP1168 Reveals Lineage-Specific Gene Loss and Genome Complexity
Bibliographic record
Abstract
Cryptophytes are a diverse lineage of marine and freshwater, photosynthetic and secondarily nonphotosynthetic algae that acquired their plastids (chloroplasts) by "secondary" (i.e., eukaryote-eukaryote) endosymbiosis. Consequently, they are among the most genetically complex cells known and have four genomes: a mitochondrial, plastid, "master" nuclear, and residual nuclear genome of secondary endosymbiotic origin, the so-called "nucleomorph" genome. Sequenced nucleomorph genomes are ∼1,000-kilobase pairs (Kbp) or less in size and are comprised of three linear, compositionally biased chromosomes. Although most functionally annotated nucleomorph genes encode proteins involved in core eukaryotic processes, up to 40% of the genes in these genomes remain unidentifiable. To gain insight into the function and evolutionary fate of nucleomorph genomes, we used 454 and Illumina technologies to completely sequence the nucleomorph genome of the cryptophyte Chroomonas mesostigmatica CCMP1168. At 702.9 Kbp in size, the C. mesostigmatica nucleomorph genome is the largest and the most complex nucleomorph genome sequenced to date. Our comparative analyses reveal the existence of a highly conserved core set of genes required for maintenance of the cryptophyte nucleomorph and plastid, as well as examples of lineage-specific gene loss resulting in differential loss of typical eukaryotic functions, e.g., proteasome-mediated protein degradation, in the four cryptophyte lineages examined.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".