MétaCan
Menu
Retour à la cohorte
Enregistrement W1996694965 · doi:10.1002/prot.22299

An intact SAM‐dependent methyltransferase fold is encoded by the human endothelin‐converting enzyme‐2 gene

2008· article· en· W1996694965 sur OpenAlexafffundabout
W. Tempel, Hong Wu, Ludmila Dombrovsky, Hong Zeng, P. Loppnau, Haizhong Zhu, A.N. Plotnikov, Alexey Bochkarev

Notice bibliographique

RevueProteins Structure Function and Bioinformatics · 2008
Typearticle
Langueen
DomaineBiochemistry, Genetics and Molecular Biology
ThématiqueCancer-related gene regulation
Établissements canadiensStructural Genomics ConsortiumUniversity of Toronto
Organismes subventionnairesDivision of Materials ResearchNational Institute of General Medical SciencesNational Cancer InstituteOffice of ScienceBasic Energy SciencesKnut och Alice Wallenbergs StiftelseKarolinska InstitutetStiftelsen för Strategisk ForskningNational Center for Research ResourcesOntario Genomics InstituteWellcome TrustNational Institutes of HealthOntario GenomicsGenome CanadaOntario Innovation TrustGlaxoSmithKlineArgonne National LaboratoryU.S. Department of EnergyNational Science Foundation
Mots-clésGene isoformBiologyRefSeqPeptide sequenceMolecular biologyBiochemistryTransmembrane domainExonN-terminusC-terminusGeneAmino acid

Résumé

récupéré en direct d'OpenAlex

A recent survey of protein expression patterns in patients with Alzheimer's disease (AD) has identified ece2 (chromosome: 3; Locations: 3q27.1) as the most significantly downregulated gene within the tested group.1ece2 encodes endothelin-converting enzyme ECE2, a metalloprotease with a role in neuropeptide processing.2 Deficiency in the highly homologous ECE1 has earlier been linked3 to increased levels of AD-related β-amyloid peptide in mice,4 consistent with a role for ECE in the degradation of that peptide.5 Initially, ECE2 was presumed to resemble ECE1,6, 7 in that it comprises a single transmembrane region of ∼20 residues flanked by a small amino-terminal cytosolic segment and a carboxy-terminal lumenar peptidase domain.8 The carboxy-terminal domain has significant sequence similarity to both neutral endopeptidase, for which an X-ray structure has been determined,9 and Kell blood group protein.10 After their initial discovery, multiple isoforms of ECE111 and ECE212 were discovered, generated by alternative splicing of multiple exons. The originally described ece2 transcript, RefSeq NM_174046,13 contains the amino-terminal cytosolic portion followed by the transmembrane region and peptidase domain (Fig. 1, isoform B). Another ece2 transcript, available from the Mammalian Gene Collection under MGC2408 (Fig. 1, isoform C), RefSeq accession NM_032331, is predicted to be translated into a 255 residue peptide with low but detectable sequence similarity to known S-adenosyl-L-methionine (SAM)-dependent methyltransferases (SAM-MTs), such as the hypothetical protein TT1324 from Thermus thermophilis, PDB14 code 2GS9,15 which shares 30% amino acid sequence identity with ECE2 over 138 residues of the sequence. Intriguingly, another “elongated” ece2 transcript (Fig. 1, isoform A) (RefSeq NM_014693) contains an amino-terminal portion of the putative SAM-MT domain, the transmembrane domain, and the protease domain. This suggests the possibility for coexistence of the putative SAM-MT and protease domains in a single polypeptide and their transmembrane interplay. Chromosomal organization of the ece2 gene and its splicing isoforms A, B, and C. Noncoding (either nontranscribed or nontranslated) regions of ece2 are shown in purple. Each box signifies one exon, with the exception of the striped grey box, which symbolizes 16 exons. Numbers in italics at the top show nucleotide positions on ece2. Numbers in bold are approximate positions on the given ECE2 isoform's amino acid sequence. Isoform B excludes all components of the SAM-MT fold, whereas isoform A lacks the C-terminal canonical SAM-MT substrate-binding subdomain. Isoform C exhibits all components of the SAM-MT fold. The diagram is not to scale. Although sequence conservation across the SAM-MT family is weak, the structural fold is highly conserved.16, 17 The most conserved part of this fold is the SAM-binding subdomain, which is shared between MGC2408 and hypothetical protein TT1324. Typically, the SAM-binding subdomain is flanked by a variable N-terminal extension and, at the C-terminus, by a substrate-binding subdomain, which varies enormously in size but preserves a conserved topology with three antiparallel β-strands. The “elongated” transcript of ece2 lacks this substrate-binding subdomain. To test the hypothesis that the 255 residue ece2 gene product MGC2408 represents a complete SAM-MT fold, we have determined a crystal structure of this protein in the presence of SAH. Cloning, protein expression, and purification: gene constructs encoding amino acid residues 19–231 of MGC2408 were amplified by PCR from the Mammalian Gene Collection18 clone (accession code gi:14150112) and subcloned into a modified pET-28a vector (gi:145307000). Subsequent sequencing of the final expression vector indicated that it has a single mutation that translates into the R100C substitution in the protein. The corresponding constructs were transformed into Escherichia coli BL21 (DE3) codon+ RIL (Stratagene) and the cells grown either in Terrific Broth (native protein) or in M9 minimal medium (Se-Met substituted protein) in the presence of 50 μg/mL of kanamycin at 37°C to an OD600 of 1.5. Cells were then induced by addition of isopropyl-1-thio-D-galactopyranoside, final concentration 1 mM, in the presence of 50 mg/L of SeMet for the SeMet substituted protein, and incubated overnight at 15°C. Cells were harvested by centrifugation at 12,227g. The cell pellets were frozen in liquid nitrogen and stored at −80°C. For purification, 11 g of cell paste were thawed and resuspended in 110 mL lysis buffer (1 XPBS, 0.25 M NaCl, 5 mM imidazole, 2 mM β-mercaptoethanol, 5% glycerol) containing 1 mM phenylmethyl sulfonyl fluoride. The cells were lysed by passing through Microfluidizer (Microfluidics Corp.) at 20,000 psi. The lysate was clarified by centrifugation and then loaded on a 5-mL HiTrap Chelating column (Amersham Biosciences), charged with Ni2+. The column was washed with 10CV of 20 mM Tris-HCl buffer, pH 8.0, containing 250 mM NaCl and 50 mM imidazole, and the protein eluted with elution buffer (20 mM Tris-HCl, pH 8.0, 250 mM NaCl, 250 mM imidazole). The protein was loaded on Superdex200 column (26 × 60) (Amersham Biosciences), equilibrated with 20 mM Tris-HCl buffer, pH 8.0, and 150 mM NaCl, at flow rate 4 mL/min. Human thrombin (Sigma) was added to the combined fractions containing MGC2408 and incubated overnight at 4°C. The protein was further purified to homogeneity by ion-exchange chromatography on Source 30Q column (10 × 10) (Amersham Biosciences), equilibrated with buffer 20 mM Tris-HCl, pH 8.0, and eluted with linear gradient of NaCl up to 500 mM concentration (20 CV). Purification yield was 4.5 mg of the protein per 1 L of culture. Crystallization: purified MGC2408 was mixed with S-adenosyl-L-homocysteine (SAH) (Sigma) at 1:5 molar ratio of protein:SAH. The sample was then crystallized using the hanging drop vapor diffusion method at 20°C by mixing 1.5 μL of the protein solution with 1.5 μL of the reservoir solution containing 28% PEG3350, 0.1M (NH4)2SO4, 0.1M BisTris, pH 5.5. Diffraction experiments, structure solution, model refinement, and analysis: diffraction data for native MGC2408 were measured at APS beamline 23ID (Argonne National Laboratory) and scaled to 1.3 Å resolution. For SeMet crystals, a 1.65 Å data set was collected at CHESS beam line A1 (Cornell University). In both cases, data were processed with HKL2000.19 Selenium substructure solution and initial phasing were performed with SHELXD20 and SHELXE,21 respectively, using the SAD22 method. ARP/wARP23 was used for automated model building. The model was refined by iteration of interactive rebuilding, restrained refinement and validation using COOT,24 REFMAC,25 and MOLPROBITY,26 respectively, and deposited27 to the Protein Data Bank14 (PDB ID: 2PXX). Data collection and processing details are listed in Table 1. Several fragments of the MGC2408 coding region were amplified and overexpressed to identify a version of the protein that would be amenable to crystallization. The construct encoding the amino-terminal Gly-Ser (remaining after proteolytic removal of the purification tag; residues 17 and 18) followed by residues 19 through 231 of the target protein (referred to as MGC2408 below in the text) produced crystals that belonged to space group P21 and diffracted to high resolution. Attempts to solve the structure by molecular replacement were unsuccessful. A selenomethionyl derivative28 was used in conjunction with SAD phasing to obtain an initial electron density map for the structure. The current model was refined against native diffraction data to 1.3 Å resolution to an Rcryst of 21.0% (Rfree 22.6%, Ref. 29). Bond lengths and angles deviate from the dictionary30, 31 values by RMSDs of 0.016 Å and 1.3°, respectively. PROCHECK32 shows 91.3% of amino acid residues in most favored and the remaining 8.7% in additional allowed regions of the Ramachandran plot.33 The final model comprises all the residues in the construct, except the two disordered carboxy-terminal residues (residues 230 and 231). The structural analysis revealed that the protein possesses the archetypal SAM-MT fold: a 3214576-ordered β-sheet sandwiched by helices with β-strand 7 oriented in antiparallel to the other strands (Fig. 2). Difference density that matched SAH was found early in the refinement process (Fig. 3). Amino acid residues 17 through 160 of our model, encoded by exons I and II of ece2, form the SAM-binding subdomain (Fig. 1). The C-terminal part of the protein (corresponding to exon III; amino acid residues 161 through 229) is folded into a typical SAM-MT substrate binding domain, which is centered around three antiparallel β-strands. A DALI34 search of the PDB revealed hypothetical protein PH0226 from Pyrococcus horikoshii OT3 (PDB ID 1VE3)35 as the closest structural homolog of MGC2408. Overall structure of MGC2408. Colors correspond to Figure 1. SAH is shown in stick representation. View of the SAH binding site highlighting key protein-ligand interactions. The model phased Fo-Fc omit map for SAH, contoured around the ligand at 3 σ and using data to 1.3 Å resolution, is shown here in purple. In contrast to the fold, which is well-conserved across the SAM-MT family, the residues composing the SAM/SAH-binding site in this family vary considerably.17 In the case of MGC2408, the aromatic side chain of Tyr-89 is aligned with the SAH adenine system through π-π-stacking. The six-amino group of the adenine ring is bound by the carboxylate of Asp-113. Hydroxyl groups in the 2′- and 3′-positions are bidentally bound by the carboxylate side chain of Asp-88. In addition, the 3′-hydroxyl group also undergoes hydrogen bonding with the side chain of Tyr-30. The homocystyl carboxylate interacts with the ε-ammonium group of Lys-130 and the peptidic NH of Trp-41, whereas the homocystyl ammonium group is bound by the peptidic carbonyl oxygens of Gly-66 and Lys-130. In an attempt to identify a putative substrate for MGC2408, the DALI search was repeated using the substrate recognition subdomain (residues 161–255). The three most homologous structures identified by this search were a DNA-binding protein SSO10B (PDB ID: 1UDV; Z score of 3.9; 12% sequence identity),36 glucose-inhibited division protein B (PDB ID: 1JSX; Z score of 3.8; 12% sequence identity),37 and a substrate recognition domain of ribosomal protein L11 methyltransferase (PDB ID: 3CJT; Z score of 3.7; 20% sequence identity).38 A reliable prediction of MGC2408 substrate specificity is not possible based on the relatively low homology with MGC2408 and the wide specificity range of these search hits. Our structure indicates that despite low sequence similarity of the C-terminal subdomain of MGC2408 with any known SAM-MTs, it adopts the typical SAM-MT fold. Taken together with information about alternative splicing variants of ece2, our structure suggests that the N-terminal SAM-binding subdomain may exist as a folded protein on its own, without the C-terminal substrate binding subdomain, as in isoform A (Fig. 1). In that protein, the SAM recognition subdomain would be followed by a transmembrane helix and a carboxy-terminal metalloendopeptidase domain. This structural arrangement would be unusual; all known SAM-MTs combine SAM and substrate recognition domains in a single polypeptide. The SAM/SAH-binding subdomain may be associated with an uncharacterized function other than MT function. Another possible explanation is that the substrate binding subdomain may be encoded by another peptide. It is further possible that a presently uncharacterized ECE2 isoform includes both the entire SAM-MT transferase (from introns I–III) and peptidase domains. This research was supported by the Structural Genomics Consortium, a registered charity (number 1097737) that receives funds from the Canadian Institutes for Health Research, the Canadian Foundation for Innovation, Genome Canada through the Ontario Genomics Institute, GlaxoSmithKline, Karolinska Institutet, the Knut and Alice Wallenberg Foundation, the Ontario Innovation Trust, the Ontario Ministry for Research and Innovation, Merck & Co., Inc., the Novartis Research Foundation, the Swedish Agency for Innovation Systems, the Swedish Foundation for Strategic Research and the Wellcome Trust. Results shown in this report are derived in part from work performed at Argonne National Laboratory, GM/CA CAT at the Advanced Photon Source funded in whole or in part with Federal funds from the National Cancer Institute (Y1-CO-1020) and the National Institute of General Medical Science (Y1-GM-1104). Use of the Advanced Photon Source was supported by the U.S. Department of Energy, Basic Energy Sciences, Office of Science, under contract No. DE-AC02-06CH11357. This work is based in part upon research conducted at the Cornell High Energy Synchrotron Source (CHESS), which is supported by the National Science Foundation under award DMR 0225180, using the Macromolecular Diffraction at CHESS (MacCHESS) facility, which is supported by award RR-01646 from the National Institutes of Health, through its National Center for Research Resources.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,000
score de la tête « metaresearch » (Gemma)0,000
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Expérimental (laboratoire) · Signal consensuel: Expérimental (laboratoire)
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,046
Score d'incertitude au seuil0,675

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0000,000
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,009
Tête enseignante GPT0,226
Écart entre enseignants0,218 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeExpérimental (laboratoire)
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations8
Publié2008
Routes d'admission3
Résumé présentoui

Explorer davantage

Même revueProteins Structure Function and BioinformaticsMême sujetCancer-related gene regulationTravaux en français237 207