Phylogeny of Primary Photosynthetic Eukaryotes as Deduced from Slowly Evolving Nuclear Genes
Notice bibliographique
Résumé
Present address: School of Life Sciences, Fudan University, Shanghai 200433, China. Present address: Division of Biological Science, Graduate School of Science, Nagoya University, Furo-cho, Chikusa-ku, Nagoya-shi, Aichi 464-8602, Japan. The biodiversity of photosynthetic eukaryotes, traditionally recognized as nine algal divisions or phyla, is attributed to two kinds of endosymbiotic events involving plastids: primary endosymbiosis and secondary endosymbiosis. Therefore, the phylogenetic positions of primary photosynthetic eukaryotes are fundamental for understanding the evolution of eukaryotic cells and establishing higher taxonomic concepts of eukaryotes. Recently, Rodríguez-Ezpeleta et al. (2005) demonstrated the strong monophyly of the three groups of primary photosynthetic eukaryotes (green plants, glaucophytes, and red algae) based on 143 nuclear genes. However, they analyzed only two divisions of the secondary phototrophs belonging to the Stramenopiles–Alveolata (SA) lineage, and their 143 genes included rapidly evolving genes. Here, we reexamine the phylogeny of the primary phototrophs based on slowly evolving nuclear genes selected mainly from the data matrix of Rodríguez-Ezpeleta et al. (2005), using additional operational taxonomic units (OTUs) of a free-living, secondary phototrophic group (Haptophyta) and of Excavata (Heterolobosea and Reclinomonas) that do not belong to the SA lineage. Our phylogenetic results demonstrate the robust non-monophyly of the primary phototrophs and the basal position of red algae within the bikonts, suggesting the loss of plastids in certain eukaryotic lineages under the assumption of the single plastid primary endosymbiosis. Maximum parsimony (MP) (with 84% bootstrap values [BT]) and Bayesian inference (BI) using the WAG+I+Γmodel (with 0.99 posterior probabilities [PP]) (see Supplemental Methods) based on the 5216 × 31 matrix (see Methods) robustly resolved the red algae as the most basal lineage within the bikonts sensu Cavalier-Smith (2003) or Plantae sensu Nozaki et al. (2003) (three groups of primary photosynthetic eukaryotes, SA lineage, and the Haptophyta; fig. 1A). In the maximum likelihood (ML) analyses (using PhyML and Proml of PHYLIP; see Supplemental Methods), however, the basal position of the red algae had only weak BT (54%–55%). This weak support may have resulted from the large amount of missing data on the glaucophyte OTUs (21%–32% of 5,216 amino acid positions versus 8.3% in total [5216 × 33 matrix]) because the ML analysis excluding these two glaucophyte species (5216 × 31-2GL matrix) showed increased support (80%–90% BT) for the most basal position of the red algae within the bikonts (fig. 1B), and addition of Glaucocystis (with 32% data missing) markedly reduced the support, especially in the ML analyses (Table S2 in Supplemental Material). In the second data matrix (5216 × 33 matrix; see Methods) that includes two OTUs of Excavata, the MP analyses and BI with the WAG+I+Γmodel, with relatively high supports (with 82% BT and 0.99 PP, respectively), resolved the most basal position of the red algae plus Excavata within the bikonts (fig. 2A), whereas the ML analyses did not resolve their basal position with 50% or more BT. However, the ML calculations excluding these two glaucophyte OTUs (5216 × 33-2GL matrix) showed increased support (51%–87% BT) for the most basal position of the red algae plus Excavata within the bikonts (fig. 2B). Bayesian inference based on the CAT+Γmodel also supports the most basal position of the red algae or red algae plus Excavata within the bikonts, with 1.00 PP (5216 × 31-2GL and 5216 × 33-2GL matrices), 0.90 PP (5216 × 31 matrix), or 0.65 PP (5215 × 33 matrix). Phylogeny based on nuclear encoded protein sequences including Haptophyta, using a 5216 × 31 matrix (31 OTUs) (A) and a 5216 ×31-2GL matrix (29 OTUs; excluding glaucophytes) (B). The analysis is based on the concatenated data set of slowly evolving nuclear proteins (19 proteins; 5,216 amino acid positions). The tree has been inferred with Bayesian inference with the WAG+I+Γmodel. Posterior probabilities (PP) for all branches are 1.00 except for branches with PP <1.00 (within brackets). Numbers above the branches represent bootstrap values (BT; ≥50%) by maximum parsimony analysis (1,000 replicates). Numbers without or within parentheses below the branches represent BT ≥50% obtained with 1,000 replicates of maximum likelihood (ML) analysis with PhyML (WAG+I+Γmodel) or Proml of PHYLIP (JTT+I+Γ model with global rearrangements), respectively. Numbers in squares show BT (10,000 replicates) calculated by the RELL method in the exhaustive ML analysis (JTT-F+Γ model). Numbers in parentheses just after the species names show possession of primary (1) or secondary (2) plastids. For details, see Supplemental Material. Phylogeny based on nuclear encoded protein sequences including Haptophyta and Excavata, using the 5216 ×33 matrix (33 operational taxonomic units [OTUs]) (A) and the 5216 ×33-2GL matrix (31 OTUs; excluding glaucophytes) (B). For details, see fig. 1. The highest likelihood trees in the exhaustive ML analyses of the 5216 × 31-2GL and 5216 × 33-2GL matrices favored polyphyletic relationships for primary photosynthetic eukaryotes (Tables S3 and S4 in Supplemental Material). The most basal group within the bikonts was composed of the red algae or the red algae plus Excavata, supported with 95% or 88% BT, using the 5216 × 31-2GL or 5216 × 33-2GL matrices, respectively (figs. 1B, 2B). In the 5216 × 31-2GL matrix, the grouping of green plants with red algae was not rejected at the 5% level by the AU, KH, or WSH test (Table S3). However, this grouping was rejected at the 5% or 1% level , respectively, by the AU or KH test in the 5216 × 33-2GL matrix (table S4). In addition, all seven trees that were not rejected by both the AU and the KH test at the 5% level (Trees 1–5, 7, and 8; Table S4) resolved that the red algae or red algae plus Excavata constitute the most basal lineage within the bikonts. Based on the very conserved nuclear genes (actin, elongation factor one alpha [EF-1α], α-tubulin, and β-tubulin), the basal phylogenetic position of the red algae within the bikonts was resolved robustly (Nozaki et al. 2003; Nozaki 2005). This phylogenetic result may have arisen from the possible relaxation of the unusually high substitution rates of the α- and β-tubulin genes in eukaryotes lacking flagellae (e.g., red algae, Dictyostelium). However, excluding these two genes, our slowly evolving gene sequences still robustly resolved the basal position of the red algae or the red algae plus Excavata within the bikonts. In addition, the present data matrix including Excavata sequences strongly rejected the monophyly between green plants and red algae in the AU and KH tests. Therefore, the strong monophyly of the three groups of primary phototrophs (Rodríguez-Ezpeleta et al. 2005) may have been due to long branch attraction (LBA) between the Opisthokonta/Amoebozoa and the SA lineage based on the fast evolving genes within the 143 genes. The SA lineage consists mainly of parasites (apicomplexans) and a ciliate (Rodríguez-Ezpeleta et al. 2005), which might have high amino acid substitutions or saturation, especially in fast-evolving genes, as a result of parasitism (Musto et al. 1999; Castro, Austin, and Dowton 2002) and atypical transcription/translation (Brunk 1986; Lozupone, Knight, and Landweber 2001). Under the assumption of a single event of plastid primary endosymbiosis (Matsuzaki et al. 2004; Rodríguez-Ezpeleta et al. 2005; for an alternative viewpoint, see Stiller, Reel, and Johnson 2003), the nonmonophyly of the primary phototrophs suggested here may be explained by the ancient primary endosymbiosis and the subsequent loss of the primary plastids in the primary plastid-lacking organisms within the bikonts (Nozaki et al. 2003; Nozaki 2005). This hypothesis may also be suggested based on the presence of cyanobacterial or plant-like genes in the nuclei of the plastid-lacking bikonts (Andersson and Roger 2002; Nozaki et al. 2003; Nozaki 2005). In any case, further phylogenetic analyses including other lineages of secondary photosynthetic eukaryotes and related nonparasitic eukaryotes are needed to resolve the correct and reliable evolutionary history of the primary plastids. As multigene analyses are expected to be increasingly sensitive to LBA, improved taxon sampling and the selection of positions or genes that evolve more slowly have been suggested for resolving deep branching in phylogenies (Philippe and Laurent 1998; Philippe, Lartillot, and Brinkmann 2005). In addition, we avoided a single OTU in each of the major lineages within the phylogenetic tree. Therefore, we analyzed only 19 slowly evolving genes, and used six additional OTUs from Haptophyta (Haptophyceae and Pavlova), Excavata (Heterolobosea [Naegleria and Sawyeria] and Reclinomonas), the red alga Galdieria, and the amoeba Physarum. The 19 genes used in this study lack complete deletion of a gene in Physarum and both two-glaucophyte OTUs, and their p-distances do not exceed 0.4 in pairwise distances (based on saturation curves of the distance-correction methods [Philippe and Laurent 1998]) for any combination of OTU for each gene, except for a combination (99 amino acids) between Dictyostelium and Toxoplasma rps17 genes (p-distance = 0.40404), and a short alignment (39 amino acids) between the Cyanophora and Physarum nsf1-I genes (p-distance = 0.46154), as well as 16 combinations (p-distance ≤ 0.43443) in rpl2, rpl27 and pls3 genes related to the Excavata. Thus, two data matrices without and with the Excavata OTUs were analyzed in this study: the “5216 × 31 matrix” consisting of 5,216 amino-acid sequences (19 genes) from 31 OTUs (excluding Excavata) and the “5216 × 33 matrix” including two OTUs of Excavata. Eighteen genes were selected from the 143 genes of Rodríguez-Ezpeleta et al. (2005) (see Supplementary Material), and the remaining gene was hsp90, which has been widely used to determine the macrophylogeny of eukaryotes in other studies (e.g., Harper, Waanders, and Keeling 2005). Because α- and β-tubulin sequences might be relaxed in eukaryotes lacking flagella (e.g., red algae, Dictyostelium), and because EF-2 protein sequences might contain unusual phylogenetic information (Stiller, Riley, and Hall 2001), we did not use these three genes. The OTUs analyzed here were the same as those of Rodríguez-Ezpeleta et al. (2005), except for the six additional OTUs (see above), and the exclusion of seven OTUs: Tetrahymena having atypical transcription and translation in gene expression (Brunk 1986; Lozupone, Knight, and Landweber 2001), and six OTUs (Babesia, Hydra, Phanerochaete, Plasmodium, Theileria, and Ustilago) from the Opisthokonta and Alveolata based on their deletion in sequences and/or high substitutions. Because the glaucophyte sequences contained the large amount of missing data (21%–32% of 5,216 amino acid positions), and because such gaps seemed to reduce the phylogenetic resolution (Table S2), two data matrices excluding the Glaucophyta (5216 × 31-2GL matrix [5216 × 31 matrix excluding glaucophytes] and the 5216 × 33-2GL matrix [5216 × 33 matrix excluding glaucophytes]) were also analyzed in this study. We are grateful to Dr. N. Rodríguez-Ezpeleta (Université de Montréal, Canada), who kindly provided the alignments of the 143 nuclear proteins. Computation time was provided by the Super Computer System, Human Genome Center, Institute of Medical Science, University of Tokyo. This work was supported by a Grant-in-Aid for Creative Scientific Research (No. 16GS0304) from the Ministry of Education, Culture, Sports, Science and Technology, Japan.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».