NMR structure of protein YvyC from <i>Bacillus subtilis</i> reveals unexpected structural similarity between two PFAM families
Notice bibliographique
Résumé
Protein containing 109-residue YvyC (MW = 13 kDa) from B. subtilis (gi|580862, SwissProt/TrEMBL ID YVYC_BACSU, accession number P39737) was selected as a target of the Protein Structure Initiative 2 and assigned to the Northeast Structural Genomics consortium (NESG; http://www.nesg.org) for structure determination (NESG Target ID SR482). YvyC belongs to the Pfam1 protein family FlaG (PF03646) and is required for proper assembly of bacterial flagella. The FlaG family contains 215 members (for sequence alignment of the FlaG family seeds see Fig. S1 in the supporting information). The NMR structure of YvyC presented here is the first atomic resolution structure available for this family. Self-assembly of bacterial flagella has been studied extensively in recent decades and comprehensive reviews are available.2-4 The gene encoding YvyC is a part of the fliD operon comprising the genes yvyC, fliD, fliS, and fliT.5 This operon is located immediately downstream of the gene hag (also called fliC) encoding flagellin—the major protein required for assembly of the flagellar filaments. During filament elongation, flagellin monomers are exported through the central channel of a flagellum and oligomerize at its distal end.2, 4 FliD (also named 'hook-associated protein 2′, or HAP26) forms a pentameric cap at the distal end of a flagellum and is essential for oligomerization of flagellin.2, 4, 7 FliS and FliT serve as flagellin-specific8 and FliD-specific9, 10 chaperones, respectively. It was demonstrated that mutations in yvyC gene orthologs in P. fluorescens (flaG) and V. anguillarum (ORF3) lead to phenotypes with unusually long filaments,11, 12 but the specific role of YvyC remains unknown. YvyC was cloned, expressed and purified following standard automated protocols to produce a uniformly 13C,15N-labeled protein sample.13, 14 Briefly, the full length yvyC gene from Bacillus subtlis was cloned into a pET21d (Novagen) derivative, yielding the plasmid pSR482-21.1. The resulting construct contains eight nonnative residues at the C-terminus (LEHHHHHH) to facilitate protein purification and one residue insertion (L) following the initiation codon introduced by cloning. Escherichia Coli BL21 (DE3) pMGK cells, a rare codon enhanced strain, were transformed with pSR482-21.1, and cultured in MJ9 minimal medium containing (15NH4)2SO4 and U-13C-glucose as sole nitrogen and carbon sources. U-13C,15N YvyC was purified using an AKTA Express (GE Healthcare) based two step protocol consisting of IMAC (HisTrap HP) and gel filtration (HiLoad 26/60 Superdex 75) chromatography. The final yield of purified U-13C,15N YvyC (>97% homogeneous by SDS-PAGE; 15.0 kDa by MALDI-TOF mass spectrometry) was ∼48 mg/L. In addition, a U-15N and 5% biosynthetically directed fractionally 13C-labeled sample15 was generated for stereo-specific assignment of isopropyl methyl groups. U-13C,15N and 5%13C,U-15N YvyC were dissolved, respectively, at concentrations of ∼1.0 and 1.1 mM in 95% H2O/5% D2O (20 mM MES, 100 mM NaCl, 10 mM DTT, 5 mM CaCl2, 0.02% NaN3) at pH 6.5. An isotropic overall rotational correlation time of ∼8 ns was inferred from 15N spin relaxation times, indicating that the protein is monomeric in solution under the conditions used for these NMR studies. This conclusion was further confirmed by analytic gel-filtration in 100 mM Tris, 100 mM NaCl, 250 ppm NaN3, at pH 7.5, with detection using a combination of static light scattering and refractive index (as described in Ref12); under these conditions the sample was observed to be >97% monomeric. All NMR spectra were recorded at 25°C. Five G-matrix Fourier transform (GFT) NMR experiments16 and a simultaneous 3D 15N/13Caliphatic/13Caromatic-resolved NOESY17 spectrum (mixing time 60 ms) were acquired on a Varian INOVA 750 MHz spectrometer equipped with a conventional probe. Two-dimensional constant-time [13C, 1H]-HSQC spectra with 28 ms and 56 ms constant-time delays were recorded for the 5% biosynthetically directed fractionally 13C-labeled sample on a Varian INOVA 600 MHz spectrometer equipped with a cryogenic probe to obtain stereo-specific assignments for isopropyl groups of valines and leucines.15 Spectra were processed using the program PROSA18 and analyzed using the program CARA.19 Sequence-specific backbone (HN, Hα, N, Cα) and Hβ/Cβ resonance assignments were obtained by using (4,3)D HNNCαβCα/CαβCα(CO)NHN and (4,3)D HαβCαβ(CO)NHN along with the program AutoAssign program.20 Side-chain spin system identification was accomplished by using aliphatic16, 21 and aromatic22 (4,3)D HCCH. Assignments were obtained for 100% of backbone and side-chain chemical shifts assignable with the NMR experiments listed above (excluding N-terminal NH, Lys NH, Arg NH2, OH of Ser, Thr and Tyr, 13Cγ of Asp and Asn, 13Cδ of Glu and Gln, and aromatic 13Cγ shifts; Table I). Stereo-specific assignments were obtained for all Val and Leu methyl groups and for 40% of the β-methylene groups exhibiting nondegenerate chemical shifts (Table I). Chemical shifts were deposited in the BioMagResBank on 06/15/2006 with accession code 7170. 1H-1H upper distance limit constraints for structure calculations were obtained from NOESY (Table I). In addition, backbone dihedral angle constraints were derived from chemical shifts using the program TALOS27 for residues located in well-defined secondary structure elements (Table I). The programs CYANA28, 29 and AUTOSTRUCTURE30 were used in parallel to assign NOEs by consensus, and the remaining assignments were carrieMOVEd by interactive spectral analysis.31 The final structure calculation was performed with CYANA, and the 20 conformers with the lowest target function value were refined in an ‘explicit water bath’32 using the program CNS.33 The coordinates were deposited in the Protein Data Bank on 06/15/2006 (accession code 2HC5). A high-quality NMR structure of protein YvyC (Table I) was obtained. The structure consists of three α-helices I-III (residues 9–22, 36–52, and 86–105) and three β-strands (residues 58–65, 68–75, and 81–85) forming one antiparallel β-sheet with topology A(↑), B(↓), C(↑) [Fig. 1(b)]. The helices form a three-helix bundle, which is attached to one side of the sheet. The secondary structure elements are locally and globally welldefined. The segment comprising residues 23–35 connecting helices I and II is largely disordered [Fig. 1(a,c)], which is manifested by comparably narrow NMR lines and lack of medium- and long-range NOEs. NMR structure of YvyC. (a) Backbone trace of residues 1–104 of the 20 representative CYANA conformers after superposition of backbone N, Cα and C′ atoms of the regular secondary structure elements for minimal root-mean-square deviation (RMSD). (b) Ribbon drawing of residues 1–104 of the conformer with the lowest CYANA target function. α-helices I-III are shown in red and yellow, β-strands A-C are shown in cyan, other polypeptide segments are shown in gray and the N- and C-termini are labeled as “N” and “C”. (c) Sausage representation of backbone and best defined side chains. A spline curve was drawn through the mean positions of Cα atoms of residues 1–104 with the thickness proportional to the mean global displacement of Cα atoms in the 20 conformers superimposed in (a). α-helices I-III are shown in red, β-strands A-C are shown in cyan, other polypeptide segments are shown in gray and a superposition of 34 side chains with the lowest global displacement is shown in blue. (d) Ribbon drawing of the conformers with the lowest CYANA target function of YvyC (pale green) and YkfF (grey) after superposition of Ca atoms of residues 37–46, 60–63, 70–74, 81–85, 90–98 and 14–23, 27–30, 38–42, 48–52, 59–67, respectively. For YvyC only residues 1–104 are shown. (e) Surface representation of the conformer with the lowest CYANA target function. The structure shown on the left has the same orientation as in (a-c), whereas the one on the right is rotated by 180° about the vertical axis. Surface colors represent sequence conservation among the seed sequences of the FlaG protein family calculated with ConSurf,34 with burgundy corresponding to the highest conservation and cyan—to the highest variability. The cavities are labeled as C1, C2, and C3. (f) Same as (e), but with the colors according to the electrostatic potential. All figures were prepared with the program MOLMOL.35 A search of the CATH database using the CATHEDRAL36 server did not yield a matching fold for protein YvyC, indicating that YvyC exhibits a distinct protein architecture. Moreover, a search of the PDB database for structurally similar proteins using both DALI37 and SSM38 identified protein YkfF from E. coli (NESG target ER397, PDB code 2HJJ) as the only significant (DALI Z-score = 3.3) hit. The structurally aligned fragment comprises 56 residues (r.m.s.d. of Cα atoms = 2.9 Å; sequence identity 5%) and contains five regular secondary structure elements (the β-sheet with strands A-C and α-helices II and III, [Fig 1(d)]. YkfF is the sole structural representative of Pfam family PF06006 comprising 38 proteins of unknown function. Hence, structure comparison of proteins YvyC from B. Subtilis and YkfF from E. coli indicates a possible distant homology between the thus far unrelated PFAM families PF03646 and PF06006. Analysis of conserved surface features34 within the FlaG protein family PF03646 reveals a single cluster of highly conserved residues on the surface of YvyC [Fig. 1(e)]. These residues belong to β-strands B and C. Furthermore, a search for protein surface cavities using the ProFunc39 server revealed three clefts [Fig. 1(e,f)] with volumes of ∼1,800, ∼800, and ∼660 Å3. The calculation of the electrostatic protein surface potential calculated with the program MolMol35 shows that in protein YvyC cavity 2 is charged negatively, whereas cavities 1 and 3 exhibit a mixed charge distribution [Fig. 1(f)]. Considering (i) that residues Leu 58, Glu 75, Ile 82 in cavity 1 and Glu 83 and Pro 86 in cavity 3 are highly conserved and (ii) that functional sites on protein surfaces are mostly located in the largest cavities,40 one may conclude that these two cavities are likely to be important for the molecular function. The authors thank Dr. Hunjoong Lee and Prof. Diana Murray for helpful discussions. Additional Supporting Information may be found in the online version of this article. Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,001 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».