Decrypting the Biochemical Function of an Essential Gene from Streptococcus pneumoniae Using ThermoFluor® Technology
Bibliographic record
Abstract
The protein product of an essential gene of unknown function from Streptococcus pneumoniae was expressed and purified for screening in the ThermoFluor® affinity screening assay. This assay can detect ligand binding to proteins of unknown function. The recombinant protein was found to be in a dimeric, native-like folded state and to unfold cooperatively. ThermoFluor was used to screen the protein against a library of 3000 compounds that were specifically selected to provide information about possible biological functions. The results of this screen identified pyridoxal phosphate and pyridoxamine phosphate as equilibrium binding ligands (Kd ∼ 50 pm, Kd ∼ 2.5 μm, respectively), consistent with an enzymatic cofactor function. Several nucleotides and nucleotide sugars were also identified as ligands of this protein. Sequence comparison with two enzymes of known structure but relatively low overall sequence homology established that several key residues directly involved in pyridoxal phosphate binding were strictly conserved. Screening a collection of generic drugs and natural products identified the antifungal compound canescin A as an irreversible covalent modifier of the enzyme. Our investigation of this protein indicates that its probable biological role is that of a nucleoside diphospho-keto-sugar aminotransferase, although the preferred keto-sugar substrate remains unknown. These experiments demonstrate the utility of a generic affinity-based ligand binding technology in decrypting possible biological functions of a protein, an approach that is both independent of and complementary to existing genomic and proteomic technologies. The protein product of an essential gene of unknown function from Streptococcus pneumoniae was expressed and purified for screening in the ThermoFluor® affinity screening assay. This assay can detect ligand binding to proteins of unknown function. The recombinant protein was found to be in a dimeric, native-like folded state and to unfold cooperatively. ThermoFluor was used to screen the protein against a library of 3000 compounds that were specifically selected to provide information about possible biological functions. The results of this screen identified pyridoxal phosphate and pyridoxamine phosphate as equilibrium binding ligands (Kd ∼ 50 pm, Kd ∼ 2.5 μm, respectively), consistent with an enzymatic cofactor function. Several nucleotides and nucleotide sugars were also identified as ligands of this protein. Sequence comparison with two enzymes of known structure but relatively low overall sequence homology established that several key residues directly involved in pyridoxal phosphate binding were strictly conserved. Screening a collection of generic drugs and natural products identified the antifungal compound canescin A as an irreversible covalent modifier of the enzyme. Our investigation of this protein indicates that its probable biological role is that of a nucleoside diphospho-keto-sugar aminotransferase, although the preferred keto-sugar substrate remains unknown. These experiments demonstrate the utility of a generic affinity-based ligand binding technology in decrypting possible biological functions of a protein, an approach that is both independent of and complementary to existing genomic and proteomic technologies. It is estimated that 40–60% of prokaryotic and eukaryotic genes have unknown or tentatively assigned biological functions (1Yakunin A.F. Yee A.A. Savchenko A. Edwards A.M. Arrowsmith C.H. Curr. Opin. Chem. Biol. 2004; 8: 42-48Crossref PubMed Scopus (59) Google Scholar). The protein products encoded by these genes are potentially valuable targets for therapeutic intervention. For example, understanding the functions of previously uncharacterized bacterial proteins could lead to the development of new classes of antibiotics (2Chan P.F. Macarron R. Payne D.J. Zalacain M. Holmes D.J. Curr. Drug Targets Infect. Disord. 2002; 2: 291-308Crossref PubMed Scopus (52) Google Scholar). Such antibiotics are urgently needed for treating the growing number of bacterial strains resistant to currently available therapies. The advent of new tools for proteomics and bioinformatics has facilitated the identification of protein function, achieved through classification of functional domains, patterns of expression, and binding partners. These tools are essential for defining biological function in a cellular context. The majority of known microbial drug targets contain sites of interaction with low-molecular weight ligands, for example, cofactors and metabolites. Discovering and understanding these “drug-able” sites on new proteins is of vital importance (3Strausberg R.L. Schreiber S.L. Science. 2003; 300: 294-295Crossref PubMed Scopus (254) Google Scholar). Sequence data alone is often insufficient to describe the molecular functions of proteins and their sites of interaction with small molecules. Even when tentative assignment of molecular function is possible, biochemical characterization of the expressed protein product is essential to confirm and elaborate function. High throughput biochemical methods have been developed to address this need and take full advantage of the vast quantities of biochemical data in commercial data bases. For example, chemically reactive probes have been used to study multiple members of a family of enzymes (4Adam G.C. Sorensen E.J. Cravatt B.F. Mol. Cell Proteomics. 2002; 1: 781-790Abstract Full Text Full Text PDF PubMed Scopus (172) Google Scholar). Affinity-based screening of ligands immobilized in microarrays has been employed to identify small molecule binding partners of proteins (5MacBeath G. Koehler A.N. Schreiber S.L. J. Am. Chem. Soc. 1999; 121: 7967-7968Crossref Scopus (408) Google Scholar). Other affinity-based technologies such as surface plasmon resonance and capillary electrophoresis have also been used to study small molecule binding when protein function is unknown (6Williams C. Curr. Opin. Biotechnol. 2000; 11: 42-46Crossref PubMed Scopus (44) Google Scholar). The utility of many methods is limited by their requirement for covalent modification of the target protein and/or the use of libraries of specialized molecules that may not be representative of the vast diversity of biochemical space. For screening using any available sample of soluble expressed protein and any commonly accessible chemical library, we have developed an affinity-based screening technology, ThermoFluor® (7Pantoliano M.W. Petrella E.C. Kwasnoski J.D. Lobanov V.S. Myslik J. Graf E. Carver T. Asel E. Springer B.A. Lane P. Salemme F.R. J. Biomol. Screen. 2001; 6: 429-440Crossref PubMed Google Scholar). ThermoFluor measures the enhanced thermal stability conferred by the binding of ligands to the native state of proteins and is readily adaptable to high-throughput screening in 384-well plates. No knowledge of substrates or binding partners is required to detect binding in ThermoFluor, a feature that renders it ideal for characterizing the binding of biological ligands to proteins of unknown function. In this work, we have screened CFE97 1The abbreviations used are: CFE97, cloned for expression protein 97; Dpx, dapoxyl sulfuric acid; Me2SO, dimethyl sulfoxide; CCD, charge coupled device; AHBA, 3-amino-5-hydroxybenzoic acid; PLP, pyridoxal phosphate; PMP, pyridoxamine phosphate; PIPES, 1,4-piperazinediethanesulfonic acid. from Streptococcus pneumoniae (8Thanassi J.A. Hartman-Neumann S.L. Dougherty T.J. Dougherty B.A. Pucci M.J. Nucleic Acids Res. 2002; 30: 3152-3162Crossref PubMed Scopus (194) Google Scholar), a protein of unknown function, against a functional probe library comprising commercially available biomolecules. In an earlier study, 113 genes were shown to be essential for growth of S. pneumoniae using targeted gene disruption (8Thanassi J.A. Hartman-Neumann S.L. Dougherty T.J. Dougherty B.A. Pucci M.J. Nucleic Acids Res. 2002; 30: 3152-3162Crossref PubMed Scopus (194) Google Scholar). These genes were selected to have a significant level of homology to related genes from at least two of four other bacterial species (40% global amino acid sequence identity to genes found in Bacillus subtilis, Enterococcus faecalis, Escherichia coli,or Staphylococcus epidermis) (8Thanassi J.A. Hartman-Neumann S.L. Dougherty T.J. Dougherty B.A. Pucci M.J. Nucleic Acids Res. 2002; 30: 3152-3162Crossref PubMed Scopus (194) Google Scholar)). A number of these genes have unknown or poorly characterized functions. One of these orphan proteins, CFE97, the protein product of gene sp1837 from S. pneumoniae, was cloned, expressed, purified, and screened using ThermoFluor. Information obtained from the screening hits enabled further biochemical characterization of the target, as well as the design of biochemical assays tailored to the putative function of CFE97. The results of these experiments, in combination with comparative analyses of the protein sequence, provide a clearer picture of the possible biochemical functions of CFE97. Materials—All materials were of the highest available quality and were purchased from Sigma unless otherwise specified. Canescin A was obtained from Microsource (Gaylordsville, CT) and Dpx was obtained from Molecular Probes (Eugene, OR). Cloning, Expression, and Purification of Protein—The sp1837 gene was inserted into a pET vector with a His6 tag as described previously (8Thanassi J.A. Hartman-Neumann S.L. Dougherty T.J. Dougherty B.A. Pucci M.J. Nucleic Acids Res. 2002; 30: 3152-3162Crossref PubMed Scopus (194) Google Scholar). The plasmid was transformed into the BL21(DE3) pLys-S expression strain of E. coli (Novagen) and 2 liters of culture were grown. The cell paste was lysed using the Rannie homogenizer and the supernatant from ultracentrifugation of the lysate was filtered through a 1.2-μm filter. The filtrate was loaded onto an 80-ml Ni-MCC (Pharmacia Chelating Sepharose Fast Flow) column, and the protein was eluted with a 25–500 mm imidazole gradient. Fractions were pooled based upon protein amount and purity in SDS-PAGE. The final purification step was an Amersham Biosciences Superdex-200 (26/60) size exclusion column with a running buffer of 20 mm mm and mm The final protein was as by using a of the ThermoFluor ThermoFluor assay was as described previously (7Pantoliano M.W. Petrella E.C. Kwasnoski J.D. Lobanov V.S. Myslik J. Graf E. Carver T. Asel E. Springer B.A. Lane P. Salemme F.R. J. Biomol. Screen. 2001; 6: 429-440Crossref PubMed Google using at ThermoFluor assay was developed by was into is a in the and other The protein at 2 in 50 mm mm was into an of compound in or and were with of to For the was from to in the were for and to four were using a well was by well and the four 3000 compounds comprising the functional probe library described were screened against CFE97, at a and mm of of protein and the of equilibrium binding has been described in (7Pantoliano M.W. Petrella E.C. Kwasnoski J.D. Lobanov V.S. Myslik J. Graf E. Carver T. Asel E. Springer B.A. Lane P. Salemme F.R. J. Biomol. Screen. 2001; 6: 429-440Crossref PubMed Google Salemme F.R. M. Google Scholar). the of the in of protein can be described by the is when protein is folded is when protein is and the the of a protein with The of the is as the The equilibrium for is as by the expression PubMed Scopus Google PubMed Scopus Google Scholar), and are the and of protein at a to be The thermal stability at any can be from the of to a using that the of protein is to be The equilibrium for ligand binding at any can be from its at a by a function, and and of ligand binding at any to that at a the is to be For protein the was to the in the of for ligand the was binding were at and to ∼ and ∼ The of as was ligand in thermal stability were The for and can be to the of ligand on the protein using the is the ligand and is the protein Salemme F.R. M. Google Scholar). binding was by the of ligand on protein Salemme F.R. M. Google Scholar). was in buffer or These were with protein and in the assay and the was at compound as described In ligand binding was found to be the of ligand was using by to to an and these to by of into of the functional probe library a of molecules selected from four classes of known molecules and generic natural and compounds that are or of biological and of this library were purchased in from the of compounds that were purchased and into based upon in data of biological ligands A. C. M. C. G. Nucleic Acids Res. 2004; PubMed Google Scholar)). The library was into 384-well as in or and the were at and was use to using was to cofactor to the analyses shown in for the in pyridoxal phosphate is into pyridoxamine phosphate were from the in at using a at mm PIPES, mm 2.5 mm mm amino were by and the of was for were the of the and were using of CFE97 mm mm was characterized by several of purified protein was estimated by using protein into mm PIPES, mm 2.5 mm The that CFE97 a amount of as by and The of the a in on with a 50 equilibrium data were on a at 20 were equilibrium was as by an The data were with a species or a as described previously R. J. J. P. J. Biol. Chem. 2003; Full Text Full Text PDF PubMed Scopus Google Scholar). was and the was The of protein was using was into mm PIPES, mm 2.5 mm using a column with the protein at was loaded into a in the protein cell an buffer in the cell was as the was to at were alone or with a of or canescin A. were in buffer to into and onto a was using in were using the is a to charge by the for a and is an estimated the charge at the of the of orphan proteins, by is by to the state of the protein in It is to the expressed protein is in a state consistent with folded protein. on a protein can protein and the of ligand on it is to a characterization of the target protein to any ThermoFluor In the of an unknown protein the data may also be in of function. purified CFE97 significant by and a in the native molecular weight was when results obtained with size exclusion of the purified CFE97 protein a an of when with a of protein consistent with a of the protein molecular is based on an of CFE97 was as a this we ultracentrifugation equilibrium data for CFE97 at several of the data to a species an of that CFE97 is a at a global was obtained to the multiple data with a a equilibrium of experiments were also with CFE97 at and of these the protein as a species with a of and not The thermal stability of CFE97 was using in the and of used in ThermoFluor experiments In the of the a stability in the of consistent with specifically with protein. could be by a equilibrium a to the data was obtained a equilibrium M.J. E. J. Mol. Biol. PubMed Scopus Google Scholar), folded is in equilibrium with These are shown in with in the results of several analyses that CFE97 as a in A of and a of 3000 were used in 2 and to the of ligands on for CFE97 Dpx Dpx, the as is on protein the of a thermal using protein is this 50 Dpx, the as is on protein the of a thermal using protein is this in a new were for CFE97 and the of ligands in the functional probe library on protein thermal stability was compounds the by an amount a of these is shown in hits from a ThermoFluor screen of the functional probe library against CFE97 in a new A and in the screening was PLP, the of CFE97 by is a cofactor for several classes of enzymes and and the genes for these enzymes as as of bacterial PubMed Scopus Google Scholar). these enzymes as or other Curr. Opin. Biol. 8: PubMed Scopus Google Scholar). The binding from both of the Curr. Opin. Biol. 8: PubMed Scopus Google Scholar). Our characterization of CFE97 as a is consistent with the biochemical of the a with an in enzymes for it is a cofactor PubMed Scopus Google Scholar). The of as a ligand may with the reactive by PLP, or protein is an in to as an essential in PubMed Scopus Google Scholar), has been found to to although the and of these are not Google Scholar). In characterization of this screening we not and this lead was not of the hits of or a and other molecules. of and identification of as a ligand in the screen to to its binding to CFE97. the thermal using ThermoFluor in the and of or the in the from to an of consistent with binding of this the the by The but significant affinity of CFE97 for is consistent with a role for The of as a function of and was using with and The of ligand on is not on these Salemme F.R. M. Google but is also a binding of ligand A of the of on ligand was achieved using for to of as be two molecules. These Kd of 50 and 2.5 are shown in The in the binding of the two of the is of conferred by upon binding of The of the in the and to at the highest is consistent with covalent binding Salemme F.R. M. Google Scholar). for CFE97 with mm PLP, not an in molecular The thermal of 20 other proteins have been in the and of PLP, and of binding have been in these The binding of PMP, although that of PLP, is of a interaction with the protein at the pyridoxal binding these results that the interaction and CFE97 is and may a requirement for enzymatic function. Canescin A ligand in the ThermoFluor screen was canescin a natural product that growth by an unknown J. 2003; PubMed Scopus Google Scholar). This compound is of the in the functional probe library, such as natural products and generic that have biological for the target of interaction is not For such ThermoFluor can be used as a to identify targets for a of to other methods of target identification C. 2004; Scopus Google Scholar). The of the of canescin A on CFE97 thermal stability was obtained to a binding The a of of ligand This on is of irreversible covalent ligands, to molecules that to the well a was for and CFE97 could not both and canescin in thermal stability was when was to a of CFE97 that been with of canescin A not covalent binding of canescin A to CFE97 was by The molecular of the protein was the protein a molecular of The was to the molecular of canescin canescin A to CFE97. Our data may also be of and canescin A for a binding but this remains to be This may that the protein target for this antifungal compound is a enzyme. Sequence and of CFE97 and D.J. J. Mol. Biol. PubMed Scopus Google of the sequence data many hits with sequence identity to CFE97 The protein found was a protein from Streptococcus 1999; PubMed Google with sequence identity to CFE97. in of ThermoFluor many of these hits known or putative functions. For example, from Bacillus A. A. G. A. M. E. M. T. G. R. M. R. 2003; PubMed Scopus Google Scholar), the sequence is to CFE97. these results that CFE97 has significant sequence homology with several classes of involved in the of nucleotide A D.J. J. Mol. Biol. PubMed Scopus Google of the of the protein available from the for J. G. Nucleic Acids Res. 2000; PubMed Scopus Google relatively and on the two 3-amino-5-hydroxybenzoic acid from and M. G. 1999; PubMed Scopus Google and from and J. J. J.A. J. 2002; Full Text Full Text PDF PubMed Scopus Google Scholar)). is involved in an essential function in M. G. 1999; PubMed Scopus Google S. J. Biol. Chem. Full Text Full Text PDF PubMed Scopus Google Scholar). a key role in to an amino from to J. J. J.A. J. 2002; Full Text Full Text PDF PubMed Scopus Google Scholar). was previously to be the known of J. J. J.A. J. 2002; Full Text Full Text PDF PubMed Scopus Google Scholar). of these enzymes use as to the study, a structure is available for of these enzymes with the M. G. 1999; PubMed Scopus Google and J. J. J.A. J. 2002; Full Text Full Text PDF PubMed Scopus Google Scholar)). CFE97 sequence identity and sequence with A. and identity and with S. These relatively low of sequence identity for the residues the binding of the residues to in the structure are to the residues of the number is for a of the binding of this functional For of the residues previously M. G. 1999; PubMed Scopus Google as involved in key binding are strictly in CFE97, with the a to binding residues by the of the functional and are not strictly but in the two to and to may for of the binding also in J. J. J.A. J. 2002; Full Text Full Text PDF PubMed Scopus Google Scholar). significant binding for and M. G. 1999; PubMed Scopus Google to be in CFE97. of this binding to and to are CFE97 and the to it that binding are CFE97 and For both and this of the that a with in both and is a of sequence in the pyridoxal binding sites of these The results of this sequence and the that the binding of in ThermoFluor for and the probable of CFE97 of for enzymes as their function or a are PubMed Scopus Google Scholar). ThermoFluor experiments to detect substrates for the results of an that was to probe a of amino for In this CFE97 was with to a was in the of amino A amino the to of of the to PMP, with Screening the 20 against the protein that the in the consistent with to the at the is a substrate for and its by CFE97 the that this is an a of of R. A. 2003; PubMed Scopus Google Scholar). identified amino using the an assay was used to the for enzymatic of to The of purified CFE97 not in the as be of cofactor of of an of PubMed Scopus Google that was of the amino a in the a at consistent with of a as described previously for E. coli S. S. J. Biol. Chem. Full Text PDF PubMed Google Scholar). low CFE97 in the of and this could be to the of to of of to as a function of amino acid was for of the 20 for substrates that were using for were independent of the amino and were several of for This that of the have the of CFE97 or of from CFE97, to in a full in were for substrates in the with the highest The for as a substrate is consistent with the in the ThermoFluor assay These data the utility of ThermoFluor as a generic for ligands that the stability of an by a substrate to a product that with for substrates in the in a new In an to identify molecules that could as amino the functional probe library was screened CFE97 in the of 20 and 50 these was to the binding that the of may be equilibrium binding ligands, or may as in the to is the acid product by the of with and is in the the by the to the and as a for the other and were also identified as of and of these are of was for J. J. J.A. J. 2002; Full Text Full Text PDF PubMed Scopus Google and consistent with a role for CFE97 as a is a substrate for an found in the for of from this is not known to in two in the bacterial of from The enzymes these and as substrates P. E. C. Mol. PubMed Scopus Google Scholar). of or both of these has not been as a function for CFE97. of and to the functional probe library several sugars and were found to the amino acid to that of CFE97 of several known and/or putative the that a is the amino for this a of commercially available sugars and were screened with ThermoFluor at Several hits were found that could a preferred substrate these sugars were not in their using these molecules as amino substrates could not be The of commercially available sugars and was at data not were to detect binding of any to CFE97 using ThermoFluor. the of and significant The results for in to the homology to and also may that a or other nucleotide diphospho-keto-sugar is the preferred amino substrate for this for binding of and selected sugars to mm mm mm mm mm in a new A comparison of the sequence of CFE97 to that of and further about CFE97 function. based on of the two are and M. G. 1999; PubMed Scopus Google the the and the of that a is overall for the the is consistent with a role in the of The is in CFE97 a role for this in of to a role in substrate J. J. J.A. J. 2002; Full Text Full Text PDF PubMed Scopus Google Scholar). the amino and the of may to for this substrate J. J. J.A. J. 2002; Full Text Full Text PDF PubMed Scopus Google Scholar). is in CFE97 and may a function, to the of CFE97 for as the amino it is not to and and into molecular possible such as these can to and the of molecules to be in at further defining function. ThermoFluor, an affinity-based screening technology, was used to for ligands that CFE97, the protein product of a gene of unknown function from S. The results of this screen that CFE97 and can it as a cofactor in an enzymatic characterization of CFE97 as a is consistent with the biochemical of experiments established that CFE97 as the amino in the step of an and that it can this biochemical and with protein sequence and structure has a assignment of the putative function of CFE97 be possible using sequence comparison This of data that this is to be an involved in the of a amino
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".