MétaCan
Menu
Back to cohort
Record W2748722638 · doi:10.3389/fbioe.2017.00048

Navigating the Functional Landscape of Transcription Factors via Non-Negative Tensor Factorization Analysis of MEDLINE Abstracts

2017· article· en· W2748722638 on OpenAlexaff
Sujoy Roy, Daqing Yun, Behrouz Madahian, Michael W. Berry, Lih‐Yuan Deng, Dan Goldowitz, Ramin Homayouni

Bibliographic record

VenueFrontiers in Bioengineering and Biotechnology · 2017
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicMachine Learning in Bioinformatics
Canadian institutionsUniversity of British Columbia
FundersUniversity of Memphis
KeywordsComputational biologyComputer scienceTensor (intrinsic definition)AnnotationGeneInformation retrievalKEGGFactorizationData miningTranscription factorBioinformaticsBiologyMathematicsGeneticsArtificial intelligenceGene expressionAlgorithmTranscriptomePure mathematics

Abstract

fetched live from OpenAlex

In this study, we developed and evaluated a novel text-mining approach, using non-negative tensor factorization (NTF), to simultaneously extract and functionally annotate transcriptional modules consisting of sets of genes, transcription factors (TFs), and terms from MEDLINE abstracts. A sparse 3-mode term × gene × TF tensor was constructed that contained weighted frequencies of 106,895 terms in 26,781 abstracts shared among 7,695 genes and 994 TFs. The tensor was decomposed into sub-tensors using non-negative tensor factorization (NTF) across 16 different approximation ranks. Dominant entries of each of 2,861 sub-tensors were extracted to form term-gene-TF annotated transcriptional modules (ATMs). More than 94% of the ATMs were found to be enriched in at least one KEGG pathway or GO category, suggesting that the ATMs are functionally relevant. One advantage of this method is that it can discover potentially new gene-TF associations from the literature. Using a set of microarray and ChIP-Seq datasets as gold standard, we show that the precision of our method for predicting gene-TF associations is significantly higher than chance. In addition, we demonstrate that the terms in each ATM can be used to suggest new GO classifications to genes and TFs. Taken together, our results indicate that NTF is useful for simultaneous extraction and functional annotation of transcriptional regulatory networks from unstructured text, as well as for literature based discovery. A web tool called Transcriptional Regulatory Modules Extracted from Literature (TREMEL), available at http://binf1.memphis.edu/tremel, was built to enable browsing and searching of ATMs.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.008
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.010
Threshold uncertainty score0.013

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.008
Meta-epidemiology (narrow)0.0020.000
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0100.006
Science and technology studies0.0000.000
Scholarly communication0.0010.002
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.006
GPT teacher head0.228
Teacher spread0.221 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations11
Published2017
Admission routes1
Has abstractyes

Explore more

Same venueFrontiers in Bioengineering and BiotechnologySame topicMachine Learning in BioinformaticsFrench-language works237,207