MétaCan
Menu
Back to cohort
Record W2000335716 · doi:10.1002/prot.22459

NMR structure of protein YvyC from <i>Bacillus subtilis</i> reveals unexpected structural similarity between two PFAM families

2009· article· en· W2000335716 on OpenAlexaff
Alexander Eletsky, Dinesh K. Sukumaran, Rong Xiao, Thomas Acton, Burkhard Rost, G.T. Montelione, Thomas Szyperski

Bibliographic record

VenueProteins Structure Function and Bioinformatics · 2009
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicBacterial Genetics and Biotechnology
Canadian institutionsStructural Genomics Consortium
FundersNational Institute of General Medical SciencesUniversity at BuffaloNational Institutes of Health
KeywordsFlagellinFlagellumBacillus subtilisOperonBiologyStructural genomicsGeneticsGeneProtein familySequence alignmentAccession number (library science)Computational biologyProtein structurePeptide sequenceBiochemistryGenBankMutantBacteria

Abstract

fetched live from OpenAlex

Protein containing 109-residue YvyC (MW = 13 kDa) from B. subtilis (gi|580862, SwissProt/TrEMBL ID YVYC_BACSU, accession number P39737) was selected as a target of the Protein Structure Initiative 2 and assigned to the Northeast Structural Genomics consortium (NESG; http://www.nesg.org) for structure determination (NESG Target ID SR482). YvyC belongs to the Pfam1 protein family FlaG (PF03646) and is required for proper assembly of bacterial flagella. The FlaG family contains 215 members (for sequence alignment of the FlaG family seeds see Fig. S1 in the supporting information). The NMR structure of YvyC presented here is the first atomic resolution structure available for this family. Self-assembly of bacterial flagella has been studied extensively in recent decades and comprehensive reviews are available.2-4 The gene encoding YvyC is a part of the fliD operon comprising the genes yvyC, fliD, fliS, and fliT.5 This operon is located immediately downstream of the gene hag (also called fliC) encoding flagellin—the major protein required for assembly of the flagellar filaments. During filament elongation, flagellin monomers are exported through the central channel of a flagellum and oligomerize at its distal end.2, 4 FliD (also named 'hook-associated protein 2′, or HAP26) forms a pentameric cap at the distal end of a flagellum and is essential for oligomerization of flagellin.2, 4, 7 FliS and FliT serve as flagellin-specific8 and FliD-specific9, 10 chaperones, respectively. It was demonstrated that mutations in yvyC gene orthologs in P. fluorescens (flaG) and V. anguillarum (ORF3) lead to phenotypes with unusually long filaments,11, 12 but the specific role of YvyC remains unknown. YvyC was cloned, expressed and purified following standard automated protocols to produce a uniformly 13C,15N-labeled protein sample.13, 14 Briefly, the full length yvyC gene from Bacillus subtlis was cloned into a pET21d (Novagen) derivative, yielding the plasmid pSR482-21.1. The resulting construct contains eight nonnative residues at the C-terminus (LEHHHHHH) to facilitate protein purification and one residue insertion (L) following the initiation codon introduced by cloning. Escherichia Coli BL21 (DE3) pMGK cells, a rare codon enhanced strain, were transformed with pSR482-21.1, and cultured in MJ9 minimal medium containing (15NH4)2SO4 and U-13C-glucose as sole nitrogen and carbon sources. U-13C,15N YvyC was purified using an AKTA Express (GE Healthcare) based two step protocol consisting of IMAC (HisTrap HP) and gel filtration (HiLoad 26/60 Superdex 75) chromatography. The final yield of purified U-13C,15N YvyC (>97% homogeneous by SDS-PAGE; 15.0 kDa by MALDI-TOF mass spectrometry) was ∼48 mg/L. In addition, a U-15N and 5% biosynthetically directed fractionally 13C-labeled sample15 was generated for stereo-specific assignment of isopropyl methyl groups. U-13C,15N and 5%13C,U-15N YvyC were dissolved, respectively, at concentrations of ∼1.0 and 1.1 mM in 95% H2O/5% D2O (20 mM MES, 100 mM NaCl, 10 mM DTT, 5 mM CaCl2, 0.02% NaN3) at pH 6.5. An isotropic overall rotational correlation time of ∼8 ns was inferred from 15N spin relaxation times, indicating that the protein is monomeric in solution under the conditions used for these NMR studies. This conclusion was further confirmed by analytic gel-filtration in 100 mM Tris, 100 mM NaCl, 250 ppm NaN3, at pH 7.5, with detection using a combination of static light scattering and refractive index (as described in Ref12); under these conditions the sample was observed to be >97% monomeric. All NMR spectra were recorded at 25°C. Five G-matrix Fourier transform (GFT) NMR experiments16 and a simultaneous 3D 15N/13Caliphatic/13Caromatic-resolved NOESY17 spectrum (mixing time 60 ms) were acquired on a Varian INOVA 750 MHz spectrometer equipped with a conventional probe. Two-dimensional constant-time [13C, 1H]-HSQC spectra with 28 ms and 56 ms constant-time delays were recorded for the 5% biosynthetically directed fractionally 13C-labeled sample on a Varian INOVA 600 MHz spectrometer equipped with a cryogenic probe to obtain stereo-specific assignments for isopropyl groups of valines and leucines.15 Spectra were processed using the program PROSA18 and analyzed using the program CARA.19 Sequence-specific backbone (HN, Hα, N, Cα) and Hβ/Cβ resonance assignments were obtained by using (4,3)D HNNCαβCα/CαβCα(CO)NHN and (4,3)D HαβCαβ(CO)NHN along with the program AutoAssign program.20 Side-chain spin system identification was accomplished by using aliphatic16, 21 and aromatic22 (4,3)D HCCH. Assignments were obtained for 100% of backbone and side-chain chemical shifts assignable with the NMR experiments listed above (excluding N-terminal NH, Lys NH, Arg NH2, OH of Ser, Thr and Tyr, 13Cγ of Asp and Asn, 13Cδ of Glu and Gln, and aromatic 13Cγ shifts; Table I). Stereo-specific assignments were obtained for all Val and Leu methyl groups and for 40% of the β-methylene groups exhibiting nondegenerate chemical shifts (Table I). Chemical shifts were deposited in the BioMagResBank on 06/15/2006 with accession code 7170. 1H-1H upper distance limit constraints for structure calculations were obtained from NOESY (Table I). In addition, backbone dihedral angle constraints were derived from chemical shifts using the program TALOS27 for residues located in well-defined secondary structure elements (Table I). The programs CYANA28, 29 and AUTOSTRUCTURE30 were used in parallel to assign NOEs by consensus, and the remaining assignments were carrieMOVEd by interactive spectral analysis.31 The final structure calculation was performed with CYANA, and the 20 conformers with the lowest target function value were refined in an ‘explicit water bath’32 using the program CNS.33 The coordinates were deposited in the Protein Data Bank on 06/15/2006 (accession code 2HC5). A high-quality NMR structure of protein YvyC (Table I) was obtained. The structure consists of three α-helices I-III (residues 9–22, 36–52, and 86–105) and three β-strands (residues 58–65, 68–75, and 81–85) forming one antiparallel β-sheet with topology A(↑), B(↓), C(↑) [Fig. 1(b)]. The helices form a three-helix bundle, which is attached to one side of the sheet. The secondary structure elements are locally and globally welldefined. The segment comprising residues 23–35 connecting helices I and II is largely disordered [Fig. 1(a,c)], which is manifested by comparably narrow NMR lines and lack of medium- and long-range NOEs. NMR structure of YvyC. (a) Backbone trace of residues 1–104 of the 20 representative CYANA conformers after superposition of backbone N, Cα and C′ atoms of the regular secondary structure elements for minimal root-mean-square deviation (RMSD). (b) Ribbon drawing of residues 1–104 of the conformer with the lowest CYANA target function. α-helices I-III are shown in red and yellow, β-strands A-C are shown in cyan, other polypeptide segments are shown in gray and the N- and C-termini are labeled as “N” and “C”. (c) Sausage representation of backbone and best defined side chains. A spline curve was drawn through the mean positions of Cα atoms of residues 1–104 with the thickness proportional to the mean global displacement of Cα atoms in the 20 conformers superimposed in (a). α-helices I-III are shown in red, β-strands A-C are shown in cyan, other polypeptide segments are shown in gray and a superposition of 34 side chains with the lowest global displacement is shown in blue. (d) Ribbon drawing of the conformers with the lowest CYANA target function of YvyC (pale green) and YkfF (grey) after superposition of Ca atoms of residues 37–46, 60–63, 70–74, 81–85, 90–98 and 14–23, 27–30, 38–42, 48–52, 59–67, respectively. For YvyC only residues 1–104 are shown. (e) Surface representation of the conformer with the lowest CYANA target function. The structure shown on the left has the same orientation as in (a-c), whereas the one on the right is rotated by 180° about the vertical axis. Surface colors represent sequence conservation among the seed sequences of the FlaG protein family calculated with ConSurf,34 with burgundy corresponding to the highest conservation and cyan—to the highest variability. The cavities are labeled as C1, C2, and C3. (f) Same as (e), but with the colors according to the electrostatic potential. All figures were prepared with the program MOLMOL.35 A search of the CATH database using the CATHEDRAL36 server did not yield a matching fold for protein YvyC, indicating that YvyC exhibits a distinct protein architecture. Moreover, a search of the PDB database for structurally similar proteins using both DALI37 and SSM38 identified protein YkfF from E. coli (NESG target ER397, PDB code 2HJJ) as the only significant (DALI Z-score = 3.3) hit. The structurally aligned fragment comprises 56 residues (r.m.s.d. of Cα atoms = 2.9 Å; sequence identity 5%) and contains five regular secondary structure elements (the β-sheet with strands A-C and α-helices II and III, [Fig 1(d)]. YkfF is the sole structural representative of Pfam family PF06006 comprising 38 proteins of unknown function. Hence, structure comparison of proteins YvyC from B. Subtilis and YkfF from E. coli indicates a possible distant homology between the thus far unrelated PFAM families PF03646 and PF06006. Analysis of conserved surface features34 within the FlaG protein family PF03646 reveals a single cluster of highly conserved residues on the surface of YvyC [Fig. 1(e)]. These residues belong to β-strands B and C. Furthermore, a search for protein surface cavities using the ProFunc39 server revealed three clefts [Fig. 1(e,f)] with volumes of ∼1,800, ∼800, and ∼660 Å3. The calculation of the electrostatic protein surface potential calculated with the program MolMol35 shows that in protein YvyC cavity 2 is charged negatively, whereas cavities 1 and 3 exhibit a mixed charge distribution [Fig. 1(f)]. Considering (i) that residues Leu 58, Glu 75, Ile 82 in cavity 1 and Glu 83 and Pro 86 in cavity 3 are highly conserved and (ii) that functional sites on protein surfaces are mostly located in the largest cavities,40 one may conclude that these two cavities are likely to be important for the molecular function. The authors thank Dr. Hunjoong Lee and Prof. Diana Murray for helpful discussions. Additional Supporting Information may be found in the online version of this article. Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: Bench or experimental
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.057
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0010.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.006
GPT teacher head0.213
Teacher spread0.207 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2009
Admission routes1
Has abstractyes

Explore more

Same venueProteins Structure Function and BioinformaticsSame topicBacterial Genetics and BiotechnologyFrench-language works237,207