MétaCan
Menu
Back to cohort
Record W2037826779 · doi:10.1002/prot.21143

NMR structure of hypothetical protein TA0938 from <i>Thermoplasma acidophilum</i>

2007· article· en· W2037826779 on OpenAlexaff
Daniel Monleón, Adelinda Yee, C.H. Arrowsmith, Bernardo Celda

Bibliographic record

VenueProteins Structure Function and Bioinformatics · 2007
Typearticle
Languageen
FieldMaterials Science
TopicEnzyme Structure and Function
Canadian institutionsUniversity of TorontoOntario Institute for Cancer Research
Fundersnot available
KeywordsThermoplasma acidophilumChemistryBiochemistry

Abstract

fetched live from OpenAlex

Structural proteomics is an emerging scientific field aiming to obtain one or more representative 3D structures for every structural domain family in nature by application of high throughput structure determination techniques. These structures, and the corresponding protein production vectors and resonance assignments, will provide a valuable resource for structural and functional studies of the thousands of proteins and their homologues that are targeted by the international structural proteomics efforts.1T. acidophilum is a thermoacidophilic archaeon2 that inhabits a hot and highly acidic environment in which few organisms are viable. The genome of T. acidophilum is one of the smallest among free-living organisms.3 TA0938 is a 110-residue conserved hypothetical protein from T. acidophilum with unknown function. From BLAST4, 5 database search, TA0938 closest homologues are hypothetical proteins in Sulfolobus solfataricus (Q97VD3_SULSO, ∼50% sequential identity over 90% of total length) and in Sulfolobus tokodaii (Q96XA7_SULTO, ∼36% sequential identity over 86% of total length), both with unknown structure and function. Here, we present the solution structure of TA0938 determined by NMR spectroscopy. On the basis of this solution three-dimensional structure and the amino acid sequence analysis, TA0938 seems to be an uncharacterized protein with a novel fold and a putative Zn-binding motif. These results may be the bases for posterior studies on structure-function relationship in proteins with similar fold. A recombinant protein consisting of the full sequence of TA0938 (110 amino acids) was expressed in E. coli BL21-Gold (DE3) cells containing the pET-15b expression vector (Novagen). Cells were grown at 37°C to an OD600 of 1.0–1.2 and induced with 1 mM IPTG for overnight at 15°C. The protein was purified to homogeneity using metal affinity chromatography. The purified protein contained the complete sequence of TA0938 plus additional N-terminal histidine tag (MGSSHHHHHHSS GLVPRGSH). U-15N and U-13C, 15N samples were produced in 2X M9 media supplemented with Zinc sulphate, biotin, 15N ammonium chloride, and 13C glucose. 15N-labeled or 13Cl15N-labeled protein solution was prepared in 20 mM Tris (pH = 6.7), 100 mM NaCl, 1 mM DTT, 0.01% NaN3, 1 mM benzamidine, 95% H20/5% D2O. The concentration of the purified protein in the NMR samples ranged between 0.8 and 1.5 mM. Molecular mass and purity of the 13C, 15N protein in the NMR samples were confirmed by mass spectroscopy. NMR spectra for backbone resonance assignments were recorded at 25°C on a Bruker AVANCE DRX 500 MHz spectrometer equipped with pulse-field gradient triple-resonance probes. Additional NMR spectra for side-chains resonance assignments were recorded at 25°C on a Bruker AVANCE DMX 800 MHz spectrometer equipped with pulse-field gradient triple-resonance probes. Linear prediction to double number of points was used in the 1H, 13C, and 15N indirect dimensions to improve digital resolution. Spectra were processed using the XwinNMR 3.1 software package. No significant differences were detected between 15N-HSQC recorded over the protein with and without Zn. Sparky 3.916 and home made shell and perl scripts were used for semi automated peak picking and peak lists filtering.7 Resonance assignments of TA0938 were obtained mainly by the combination of manual and automatic techniques as specified elsewhere8 (BMRB entry 6812). Overall, ∼99% of backbone assignable protons and ∼96% of total protons were available for our study. For structure calculation purposes, 13C-NOESY-HSQC (τm of 100 and 200 ms) and 15N-NOESY-HSQC (τm of 100 ms) spectra were recorded at 800 MHz (for a review on the experiments used, see Cavanagh et al.9). NOE cross-peak assignment was obtained using a combination of manual and automatic methods. A preliminary fold was calculated on the basis of manually unambiguously assigned NOEs on the 3D NOESY spectra at 100 ms of mixing time. In this stage, NOE peaks were classified as weak, medium, and strong intensity and upper limit of 5.0, 4.0, and 3.0 Å in distance restraints were accordingly applied. NOE assignments were extended and structure was further refined by spectra/structure iterative semi-automated analysis with the NOAH module in the program DYANA10 including all spectra. Peak lists of the NOESY spectra were obtained by interactive peak picking using the “restricted peak picking” option of the program SPARKY. Backbone dihedral restraints were derived from the 1Hα and 13Cα secondary chemical shifts using TALOS.11 A summary of the final set of structural restraints used for torsion angle dynamics calculations together with other statistics for the ensemble of the 20 lowest target function values conformers is reported in Table I. The program MOLMOL12 was used to analyze the 20 energy-minimized conformers with lowest NOE violations and to calculate solvent accessible surface and electrostatic charges distributions. MOLMOL was also used to prepare drawings of the structures. The 20 lowest target function structures of TA0938 are well-converged, as shown in Figure 1. The structure of TA0938 has two clearly different parts: a region containing all the regular secondary structure elements and a bundle of loops, which contain all cysteines in the protein. The part of the protein with regular secondary structure is formed by a central β-sheet flanked by two α-helices (residues 54–62 and 101–106) at each face of the sheet. In this part of the protein, a very flexible loop is observed between residues 90 and 100. Significant broadening of peaks belonging to this region seems to support the flexibility suggested by the partial disorder observed in this loop. The molecular topology of the protein can be described as βαββα. In this topology, the bundle of loops with the cysteines is located between β1 and α1. The central β-sheet is composed by three strands (residues 3–8, 68–73, and 82–85) with a β2([darrow]), β1(↑), and β3(↑) pattern. The bundle of loops where the cysteines are located (residues 16–46) shows some degree of sequence similarity to reported Zn-binding motifs. A strong NOE between Hβ of cysteines Cys20 and Cys46 suggests that these cysteines may be in disulfide bond distance range. The chemical shifts for the Cβ of these cysteines seem to support the partial formation of this disulfide bond. NMR solution structure of TA0938. A. Stereo view of the superposition of the final 20 structures over the average NMR structure. B. Ribbon diagram depicting lowest target function NMR structure of TA0938 of Thermoplasma acidophillum (PDB accession code 2FQH). The α-helices are shown in red and yellow and β-sheets are shown in cyan. Side-chains of cysteine residues are shown in blue. Chemical shift values for cysteines Cβ in TA0938 were: Cys20, 44.14 ppm; Cys23, 32.28 ppm; Cys42, 36.02 ppm; Cys43, 34.36 ppm; and Cys46 43.10 ppm C. Sausage presentation of backbone of TA0938. Thickness of the cylindrical rod is proportional to the mean of the global displacements of the Cα atoms in the 20 DYANA best conformers. The β-strands are shown in green, the α-helices in red and cysteine residues in yellow. A homology three-dimensional structure search using DALI13 within the Protein Data Bank showed that TA0938 shares no meaningful structural similarity with protein structures reported to date. The highest z score is below the threshold of being significant (i.e. 0.8 versus a threshold of 2.0). These results imply that TA0938 would be an uncharacterized protein with a novel fold. Nevertheless, the conformation detected for the region where the cysteines are located showed some structural similarity to the E. Coli CLPX chaperone Zinc-binding domain [PDB code 1OVX, Figure 2(A)]. Sequence identity between this TA0938 segment and the Zn-binding motif of CLPX was 30%. In addition, Cβ chemical shift values for Cys23, Cys 43, and Cys44 in this loop seemed consistent with Zn-binding cysteine 13C average chemical shift values.14 A. Comparison between cysteine-enriched regions of TA0938 (in red) and E. Coli CLPX chaperone Zinc binding domain dimer (PDB code 1OVX, in blue). B. Sequence alignment of TA0938, closest homologues hypothetical proteins Q97VD3_SULSO and Q96XA7_SULTO and cysteine-rich segment from Escherichia coli CLPX chaperone Zinc binding domain. Secondary structure elements in TA0938 structure have been plotted in the sequence. Cysteines have been marked inside a rectangle. In summary, we present the solution structure of TA0938, a functionally unknown protein in Thermoplasma acidophillum. According to this structure, TA0938 would be an uncharacterized protein with a novel fold with a cysteine-rich region, which resembles some Zn-binding motifs. Structures ensemble has been deposited into the Protein Data Bank15 (PDB accession code 2FQH). The authors thank the SCSIE of the University of Valencia for providing access to the NMR facility and high performance computing facilities. The NMR Facility of the Parc Cientific the Barcelona is also gratefully acknowledged for providing access to the 800 MHz NMR spectrometer. We also thank Bruker España S.A. for technical and economic support as part of the agreement with the University of Valencia for the development of new techniques in biomolecular NMR.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Insufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: Bench or experimental
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.018
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.006
GPT teacher head0.195
Teacher spread0.189 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2007
Admission routes1
Has abstractyes

Explore more

Same venueProteins Structure Function and BioinformaticsSame topicEnzyme Structure and FunctionFrench-language works237,207