MétaCan
Menu
Retour à la cohorte
Enregistrement W2006189661 · doi:10.1002/prot.20316

Crystal structure of the conserved hypothetical protein MPN330 (GI: 1674200) from <i>Mycoplasma pneumoniae</i>

2004· article· en· W2006189661 sur OpenAlexaboutno aff
Debanu Das, N. Oganesyan, Hisao Yokota, Ramona Pufan, Rosalind Kim, Sung‐Hou Kim

Notice bibliographique

RevueProteins Structure Function and Bioinformatics · 2004
Typearticle
Langueen
DomaineImmunology and Microbiology
ThématiqueMicrobial infections and disease research
Établissements canadiensnon disponible
Organismes subventionnairesNational Institute of General Medical Sciences
Mots-clésStructural genomicsMycoplasma genitaliumBiologyMycoplasma pneumoniaeHypothetical proteinGeneProtein structureGeneticsSequence alignmentProtein familyComparative genomicsPeptide sequenceGenomicsMycoplasmaComputational biologyGenomeVirologyBiochemistry

Résumé

récupéré en direct d'OpenAlex

The Berkeley Structural Genomics Center (BSGC) pursues an integrated structural genomics approach to determine the 3D crystal structures of proteins in two minimal genomes: Mycoplasma pneumoniae (MP) and M. genitalium (MG), two related human and animal pathogens. This leads to the inference of function based on protein structure and identifying novel folds in the protein fold universe, thus adding to the repertoire of the known natural folds of proteins. Due to the increasing level of antibiotic resistance in bacteria, structural genomics of pathogenic organisms can also lead to the identification of new protein drug targets for therapeutic purposes. The MPN330 (GI: 1674200) gene from M. pneumoniae encodes a 294 amino acid protein of unknown function that is highly conserved in 3 of the 10 sequenced members of the mycoplasma family (64% sequence identity to M. genitalium MG237, 26% in M. gallisepticum, 22% in M. penetrans) and moderately conserved in several others: M. mobile, Mesoplasma florum, and Ureaplasma parvum serovar. Apart from this, the MPN330 does not have any appreciable sequence similarity to any other protein as evaluated by PSI-BLAST.1 Since this protein is conserved in this family, it is likely to be involved in some cellular function that is yet to be elucidated. The crystal structure of this protein has been determined at 2.5 Å by single-wavelength anomalous diffraction (SAD) in an effort to explore its structure–function relationships. Analysis of the 3D constitution of the protein reveals that it has 11 helices and two strands distributed into three structural domains, one of which has a novel fold. An area of the protein surface has a negatively charged region flanked by some conserved residues that could be important for its function. Protein expression and purification: The MPN330 gene was amplified by PCR using M. pneumoniae genomic DNA template and primers designed for Ligation-Independent Cloning (LIC).2 The amplified PCR product was prepared for vector insertion by purification, quantitation, and treatment with T4 DNA polymerase (New England Biolabs, Beverley, MA) in the presence of 1 mM dTTP. The prepared insert was annealed into the LIC expression vector pB2, a derivative of pET21a (Novagen, Madison, WI) that expresses the cloned gene fused with an N-terminal hexa-histidine tag and transformed into E. coli DH5α for plasmid amplification. Selenomethionyl (Se-Met) protein (3 Selenium atoms in 294 residues) was prepared by a modification of the method of Doublie3 with the addition of 0.25 M NaCl in the growth media. Cells were grown at 37°C in Escherichia coli strain B834(DE3)/pSJS12444 until O.D.600 ∼0.9, transferred to 47°C and induced with 0.5 mM isopropyl-β-D-thiogalactopyranoside. Cells were kept at 47°C for 20 min and then transferred to 20°C for overnight growth. The addition of salt and heat induction were done to induce a stress response. Cell lysis was performed in a microfluidizer (Microfluidics, Newton, MA) in the presence of 50 mM Hepes (pH 8.0), 0.1 M NaCl, 10 mM βME, 1 mM PMSF, 10 μg/ml DNAse and Roche Protease Inhibitor Cocktail Tablet (catalog number 1836145, Roche Diagnostics, USA). The lysate was then spun in a Beckman ultracentrifuge in a Ti45 rotor at 35,000 rpm for 30 min at 4°C. The 6XHis-tagged protein was affinity purified from the soluble fraction using a 5-ml HiTrap Chelating HP Column (Amersham Biosciences, Piscataway, NJ) as recommended by the manufacturer and elution was achieved with a linear gradient from 10 mM to 400 mM imidazole in 13 column volumes. Ni+2 affinity purified fractions were then purified by ion-exchange chromatography using a 5-ml Hi Trap Q-Sepharose (Amersham) column in buffer containing 50 mM Hepes (pH 8.0), 50 mM NaCl using a linear salt gradient from 50 mM to 500 mM NaCl in 10 column volumes. The final purity of the sample was determined by SDS-PAGE, monodispersity was measured by dynamic light scattering using the DynaPro 99 (Proterion Corp., Piscataway, NJ), and the molecular weight was confirmed by MALDI-TOF mass spectrometry using the Voyager DE (Applied Biosystems, Foster City, CA). Crystallization and X-ray diffraction data collection: Single crystals measuring 50 × 30 × 20 μm were obtained in a 0.4 μl sitting drop vapor diffusion crystallization experiment in 1.5 M Li2SO4 and 0.1 M Tris-HCl (pH 8.5) (condition 72 of the High Throughput Salt Screen, Hampton Research, Aliso Viejo, CA) that was set up using a Hydra Crystallization Robot (Robbins Scientific, Sunnyvale, CA). After freezing a single crystal from the 0.4 μl sitting drop at 100 K using 15% glycerol as cryoprotectant, a single wavelength anomalous diffraction (SAD) data set was collected at the Macromolecular Crystallography synchrotron beamline 5.0.2 (Advanced Light Source, Lawrence Berkeley National Laboratory, Berkeley, CA) equipped with a Quantum 4 CCD detector (Area Detector System Corporation, Poway, CA) placed 250 mm from the sample. The X-ray diffraction data were integrated and scaled with HKL2000.5 The space group was determined to be I4122 with the crystal cell dimensions of 82.84, 82.84, and 225.04 Å with one protein molecule in the asymmetric unit of the unit cell. Structure determination: The heavy-atom substructure was solved by difference Patterson methods calculated from 20-3 Å as implemented in SOLVE6 and verified with PHENIX.7 With statistical density modification and phase extension to 2.5 Å as provided in RESOLVE,8 phases from two Se atoms in each molecule resulted in an interpretable electron density map in which ∼70% of the Cα backbone was built in automatically using RESOLVE_BUILD.9 Several rounds of iterative model building in O10 guided by experimental, 2Fo−Fc and Fo−Fc electron density maps and crystallographic refinement using CNS11 including simulated annealing, positional and B-factor refinement allowed the rest of the model to be built. Addition of solvent molecules and a final round of refinement resulted in a final R = 24.4% and Rfree = 29.8%. The final model is comprised of residues 5–290 with 84.2% of the residues in the core regions of the Ramachandran plot, 14.0% in additionally favorable regions and 1.8% in generously allowed regions as evaluated by PROCHECK.12 The background in the Fo−Fc electron density map was noisier in the C-terminal half of the molecule which may reflect some local disorder and may partially account for the slightly high R-factors. The difference density map also shows a peak near the surface exposed residue 21 that may be a metal ion which has not been identified. The crystal structure shows that the protein has overall dimensions of approximately 55 × 47 × 25 Å and is mainly alpha helical with 11 helices and 2 strands (Fig. 1) that can be separated into three structural domains. The N- and C-terminal domains each have four helices but with different relative orientations and sizes of the helices. This results in the domains resembling one another but they are not appreciably similar in structure and cannot be superimposed. These helical bundles in the two domains point in different directions with respect to the plane of the central domain. In a comparison with other protein structures using DALI,13 the top match found (Z-score = 5.9, RMSD of 3.1 Å) is not significantly similar over the whole length of the protein to any other protein structure in the Protein Data Bank implying that this is a new structure. However, better structural matches for two of the three domains can be found based on comparisons of the individual domains using DALI and CE.14 The N-terminal portion of MPN330 is most similar to invertase inhibitor (Z-score = 6.2, PDB code 1RJ115) and superimposes with an RMSD of 3.0 Å for Cα atoms of residues 20–102 from MPN330 and 41–147 from invertase inhibitor [Fig. 2(a)]. The sizes of the helices are not the same, being shorter in the MPN330, but their relative orientations within the domain are similar. The C-terminal portion of MPN330 is similar to the N-terminal portion of a cyclin homolog (Z-score = 6.8, PDB code 1BU216) with an RMSD of 3.1 Å including Cα atoms for amino acids 199–284 from MPN330 and 47–148 of the cyclin homolog [Fig. 2(b)]. The middle domain (residues 107–198) is comprised of three helices in sequence followed by two strands and does not appreciably resemble any other known structure (Z-score = 4.6, RMSD of 3.7 Å), thus suggesting a novel fold. Despite the structural similarity of these two domains, the residues that are functionally important in the invertase inhibitor and cyclin homolog are not conserved in this protein nor are they substituted with a different set of conserved amino acids. This suggests that this protein uses similar structural features to perform a different function. Since both invertase inhibitor and cyclin are proteins that bind to other proteins, it is possible that this protein is also some kind of regulatory protein or that it binds with specific ligands. There is also some evidence of this from an analysis of the molecular surface properties as described below. Ribbon rendering with PyMOL18 of the crystal structure of the conserved hypothetical protein MPN330 (GI: 1674200) from M. pneumoniae showing the three structural domains including residues 5–106 (green), 107–198 (blue) and 199–290 (purple). The first and the third domains have four helices each and the middle domain has three helices and two strands and represents a novel fold. a: Superposition of the N-terminal domain (5–106) from the MPN330 protein (green) with residues 1–147 (white) of the invertase inhibitor (PDB code: 1RJ1) and b: the C-terminal domain (199–290) (purple) of MPN330 with residues 47-148 (white) from the cyclin homolog (PDB code: 1BU2) showing their structural similarity. The potential energy surface of the protein rendered with GRASP17 shows that one side of the protein has a sizeable negatively charged patch that has at its center a cleft-like feature (Fig. 3). Although this protein is annotated as a hypothetical protein with unknown function, it is conserved in most of the mycoplasma family with sequence identities ranging from 20% to 64%, thereby suggesting some functional role. The sequence alignment with the other family members with the highest sequence identity (M. genitalium, M. gallisepticum, and M. penetrans) shows that many of the residues are conserved and are distributed over all three domains (Fig. 4). Some of these conserved residues are not surface-exposed and are buried in the hydrophobic core, thereby used to maintain the structural integrity of the protein. However, many of the conserved residues (F8, E69, E70, Y75, K220, E226, E230, and Y258) are present on the surface surrounding the negatively charged patch described above [Fig. 5(a)] or in the cleft. It is interesting that these conserved residues belong to the N- and C-terminal domains that come close spatially to form this cleft. It is possible that this protein functions as a specific regulator of other protein functions or in ligand binding. This can occur by protein docking with an interacting partner because of electrostatic interactions provided by the negatively charged surface and specificity of recognition being achieved by the conserved residues that fence this charged area. The negative potential in the cleft may also be a site for specific ligand binding. Some other conserved residues (E194, Q186, F188, Y206, and Y239) are solvent-exposed on another face of the protein [Fig. 5(b)] and may modulate yet another facet of its function. The presence of a number of aromatic residues on the protein surface may indicate their role in protein–protein interactions. The protein was subjected to an array of functional assays for enzymatic activity including phosphatase assay with pNPP as substrate; phosphodiesterase/nuclease assay (with pNP-TMP and bis-pNPP); amino acid dehydrogenase, alcohol dehydrogenase, and organic acid dehydrogenase assays; and assays for protease, oxidase, sulfatase, and esterase/lipase activity. The results showed that the protein does not have any of the above activities. Electrostatic surface potential using GRASP17 of the MPN330 molecule showing a negatively charged cleft on one side of the protein. Multiple sequence alignment of the M. pneumoniae protein MPN330 with its nearest homologs from M. genitalium, M. gallisepticum, and M. penetrans showing the secondary structure elements and the pattern of conserved residues. The sequence alignment was done using CLUSTALW19 and the figure generated with ESPRIPT.20 a: Conserved residues around the negatively charged cleft as shown in Figure 3. Residues E69, E70, E226 and E230 contribute to the negative patch whereas residues F8, Y75, K220, and Y258 flank the cleft. These may be involved in ligand binding or protein–protein interactions and b: conservation of residues Q186, F188, E194, Y206 and Y239 on another side of the protein (90° from the view above) suggesting their possible role in another aspect of function. The crystal structure has been deposited in the Protein Data Bank with accession number 1TD6. We thank Dr. David King for mass spectrometric analysis of the protein, Barbara Gold for cloning, Marlene Henriquez, Irina Ankoudinova, and Bruno Martinez for expression studies and cell paste preparation, Dr. John-Marc Chandonia for bioinformatics search of the gene, Dr. Vaheh Oganesyan, for help in data collection, Dr. Paul Adams for computational crystallography support, Dr. Alexander Iakounine of the University of Toronto for functional assays, and Dr. Christine Trame for user support at beamline 5.0.2 at the Advanced Light Source. The Advanced Light Source is supported by the Director, Office of Science, Office of Basic Energy Sciences, Materials Sciences Division, of the U.S. Department of Energy under Contract No. DE-AC03-76SF00098 at Lawrence Berkeley National Laboratory.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,000
score de la tête « metaresearch » (Gemma)0,000
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesCharge utile insuffisante (le modèle a refusé de juger)
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Expérimental (laboratoire) · Signal consensuel: Expérimental (laboratoire)
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,033
Score d'incertitude au seuil1,000

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0000,000
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,001
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,001
Charge utile insuffisante (le modèle a refusé de juger)0,0010,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,008
Tête enseignante GPT0,202
Écart entre enseignants0,194 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Devis d'étudeExpérimental (laboratoire)
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations4
Publié2004
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueProteins Structure Function and BioinformaticsMême sujetMicrobial infections and disease researchTravaux en français237 207