NMR structure of hypothetical protein TA0938 from <i>Thermoplasma acidophilum</i>
Bibliographic record
Abstract
Structural proteomics is an emerging scientific field aiming to obtain one or more representative 3D structures for every structural domain family in nature by application of high throughput structure determination techniques. These structures, and the corresponding protein production vectors and resonance assignments, will provide a valuable resource for structural and functional studies of the thousands of proteins and their homologues that are targeted by the international structural proteomics efforts.1T. acidophilum is a thermoacidophilic archaeon2 that inhabits a hot and highly acidic environment in which few organisms are viable. The genome of T. acidophilum is one of the smallest among free-living organisms.3 TA0938 is a 110-residue conserved hypothetical protein from T. acidophilum with unknown function. From BLAST4, 5 database search, TA0938 closest homologues are hypothetical proteins in Sulfolobus solfataricus (Q97VD3_SULSO, ∼50% sequential identity over 90% of total length) and in Sulfolobus tokodaii (Q96XA7_SULTO, ∼36% sequential identity over 86% of total length), both with unknown structure and function. Here, we present the solution structure of TA0938 determined by NMR spectroscopy. On the basis of this solution three-dimensional structure and the amino acid sequence analysis, TA0938 seems to be an uncharacterized protein with a novel fold and a putative Zn-binding motif. These results may be the bases for posterior studies on structure-function relationship in proteins with similar fold. A recombinant protein consisting of the full sequence of TA0938 (110 amino acids) was expressed in E. coli BL21-Gold (DE3) cells containing the pET-15b expression vector (Novagen). Cells were grown at 37°C to an OD600 of 1.0–1.2 and induced with 1 mM IPTG for overnight at 15°C. The protein was purified to homogeneity using metal affinity chromatography. The purified protein contained the complete sequence of TA0938 plus additional N-terminal histidine tag (MGSSHHHHHHSS GLVPRGSH). U-15N and U-13C, 15N samples were produced in 2X M9 media supplemented with Zinc sulphate, biotin, 15N ammonium chloride, and 13C glucose. 15N-labeled or 13Cl15N-labeled protein solution was prepared in 20 mM Tris (pH = 6.7), 100 mM NaCl, 1 mM DTT, 0.01% NaN3, 1 mM benzamidine, 95% H20/5% D2O. The concentration of the purified protein in the NMR samples ranged between 0.8 and 1.5 mM. Molecular mass and purity of the 13C, 15N protein in the NMR samples were confirmed by mass spectroscopy. NMR spectra for backbone resonance assignments were recorded at 25°C on a Bruker AVANCE DRX 500 MHz spectrometer equipped with pulse-field gradient triple-resonance probes. Additional NMR spectra for side-chains resonance assignments were recorded at 25°C on a Bruker AVANCE DMX 800 MHz spectrometer equipped with pulse-field gradient triple-resonance probes. Linear prediction to double number of points was used in the 1H, 13C, and 15N indirect dimensions to improve digital resolution. Spectra were processed using the XwinNMR 3.1 software package. No significant differences were detected between 15N-HSQC recorded over the protein with and without Zn. Sparky 3.916 and home made shell and perl scripts were used for semi automated peak picking and peak lists filtering.7 Resonance assignments of TA0938 were obtained mainly by the combination of manual and automatic techniques as specified elsewhere8 (BMRB entry 6812). Overall, ∼99% of backbone assignable protons and ∼96% of total protons were available for our study. For structure calculation purposes, 13C-NOESY-HSQC (τm of 100 and 200 ms) and 15N-NOESY-HSQC (τm of 100 ms) spectra were recorded at 800 MHz (for a review on the experiments used, see Cavanagh et al.9). NOE cross-peak assignment was obtained using a combination of manual and automatic methods. A preliminary fold was calculated on the basis of manually unambiguously assigned NOEs on the 3D NOESY spectra at 100 ms of mixing time. In this stage, NOE peaks were classified as weak, medium, and strong intensity and upper limit of 5.0, 4.0, and 3.0 Å in distance restraints were accordingly applied. NOE assignments were extended and structure was further refined by spectra/structure iterative semi-automated analysis with the NOAH module in the program DYANA10 including all spectra. Peak lists of the NOESY spectra were obtained by interactive peak picking using the “restricted peak picking” option of the program SPARKY. Backbone dihedral restraints were derived from the 1Hα and 13Cα secondary chemical shifts using TALOS.11 A summary of the final set of structural restraints used for torsion angle dynamics calculations together with other statistics for the ensemble of the 20 lowest target function values conformers is reported in Table I. The program MOLMOL12 was used to analyze the 20 energy-minimized conformers with lowest NOE violations and to calculate solvent accessible surface and electrostatic charges distributions. MOLMOL was also used to prepare drawings of the structures. The 20 lowest target function structures of TA0938 are well-converged, as shown in Figure 1. The structure of TA0938 has two clearly different parts: a region containing all the regular secondary structure elements and a bundle of loops, which contain all cysteines in the protein. The part of the protein with regular secondary structure is formed by a central β-sheet flanked by two α-helices (residues 54–62 and 101–106) at each face of the sheet. In this part of the protein, a very flexible loop is observed between residues 90 and 100. Significant broadening of peaks belonging to this region seems to support the flexibility suggested by the partial disorder observed in this loop. The molecular topology of the protein can be described as βαββα. In this topology, the bundle of loops with the cysteines is located between β1 and α1. The central β-sheet is composed by three strands (residues 3–8, 68–73, and 82–85) with a β2([darrow]), β1(↑), and β3(↑) pattern. The bundle of loops where the cysteines are located (residues 16–46) shows some degree of sequence similarity to reported Zn-binding motifs. A strong NOE between Hβ of cysteines Cys20 and Cys46 suggests that these cysteines may be in disulfide bond distance range. The chemical shifts for the Cβ of these cysteines seem to support the partial formation of this disulfide bond. NMR solution structure of TA0938. A. Stereo view of the superposition of the final 20 structures over the average NMR structure. B. Ribbon diagram depicting lowest target function NMR structure of TA0938 of Thermoplasma acidophillum (PDB accession code 2FQH). The α-helices are shown in red and yellow and β-sheets are shown in cyan. Side-chains of cysteine residues are shown in blue. Chemical shift values for cysteines Cβ in TA0938 were: Cys20, 44.14 ppm; Cys23, 32.28 ppm; Cys42, 36.02 ppm; Cys43, 34.36 ppm; and Cys46 43.10 ppm C. Sausage presentation of backbone of TA0938. Thickness of the cylindrical rod is proportional to the mean of the global displacements of the Cα atoms in the 20 DYANA best conformers. The β-strands are shown in green, the α-helices in red and cysteine residues in yellow. A homology three-dimensional structure search using DALI13 within the Protein Data Bank showed that TA0938 shares no meaningful structural similarity with protein structures reported to date. The highest z score is below the threshold of being significant (i.e. 0.8 versus a threshold of 2.0). These results imply that TA0938 would be an uncharacterized protein with a novel fold. Nevertheless, the conformation detected for the region where the cysteines are located showed some structural similarity to the E. Coli CLPX chaperone Zinc-binding domain [PDB code 1OVX, Figure 2(A)]. Sequence identity between this TA0938 segment and the Zn-binding motif of CLPX was 30%. In addition, Cβ chemical shift values for Cys23, Cys 43, and Cys44 in this loop seemed consistent with Zn-binding cysteine 13C average chemical shift values.14 A. Comparison between cysteine-enriched regions of TA0938 (in red) and E. Coli CLPX chaperone Zinc binding domain dimer (PDB code 1OVX, in blue). B. Sequence alignment of TA0938, closest homologues hypothetical proteins Q97VD3_SULSO and Q96XA7_SULTO and cysteine-rich segment from Escherichia coli CLPX chaperone Zinc binding domain. Secondary structure elements in TA0938 structure have been plotted in the sequence. Cysteines have been marked inside a rectangle. In summary, we present the solution structure of TA0938, a functionally unknown protein in Thermoplasma acidophillum. According to this structure, TA0938 would be an uncharacterized protein with a novel fold with a cysteine-rich region, which resembles some Zn-binding motifs. Structures ensemble has been deposited into the Protein Data Bank15 (PDB accession code 2FQH). The authors thank the SCSIE of the University of Valencia for providing access to the NMR facility and high performance computing facilities. The NMR Facility of the Parc Cientific the Barcelona is also gratefully acknowledged for providing access to the 800 MHz NMR spectrometer. We also thank Bruker España S.A. for technical and economic support as part of the agreement with the University of Valencia for the development of new techniques in biomolecular NMR.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".