MétaCan
Menu
Back to cohort
Record W4416716702 · doi:10.1101/2025.11.23.690073

A Multi-Agent Approach to Generating Context-Rich Gene Sets

2025· preprint· W4416716702 on OpenAlexaff
Ebunoluwa Makinde, Farhad Maleki, Katie Ovens

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2025
Typepreprint
Language
FieldBiochemistry, Genetics and Molecular Biology
TopicBiomedical Text Mining and Ontologies
Canadian institutionsUniversity of Calgary
Fundersnot available
KeywordsSet (abstract data type)Relevance (law)Gene nomenclaturePipeline (software)GeneSelection (genetic algorithm)Gene predictionTask (project management)Biological data

Abstract

fetched live from OpenAlex

Abstract Gene sets are collections of genes that share a common biological function, process, or component that can be used to get insight into the biological relevance of genomic data. Databases containing these gene sets aids in a wide array of analytical methods. The results of these methods, such as gene set analysis or phenotype-based gene prioritization, depend on the quality of the gene sets. Despite the extensive literature and genetic data available for constructing these databases, they often lack sufficient biological context. Current curation methods rely on labour-intensive expert manual curation from literature and datasets, as well as automated methods that are not context-aware. Therefore, there is a significant opportunity to utilize publicly available literature to bridge this gap and create more precise gene sets. With the advancement of natural language processing technologies, particularly large language models, this task can be performed more efficiently. In this work, we present a multi-agent system that utilizes the Llama 3, DeepSeek, and Qwen open-source large language models to analyze PubMed abstracts, allowing us to reconstruct gene sets in existing databases that better reflect specific biological contexts. Our approach consists of two pipelines. One verifies the inclusion of genes in a gene set by proof of evidence in the abstracts showing the association between the gene and the gene set. The second pipeline parses through the abstracts to identify genes not already included in the gene set for potential inclusion. To evaluate the proposed approach, we reconstructed a random selection of gene sets within the Human Ontology Phenotype (HPO). Our analysis shows that 149 of these gene sets have a similarity of 65.18% when compared to the original HPO gene sets, aligning well with the current HPO database. Additionally, we found an average of 3.15 new genes not included in the HPO gene sets, each supported by verified literature linking them to their respective gene sets. This highlights that our updated gene set database better reflects the current state of biological findings.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.004
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.006
Threshold uncertainty score0.013

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.004
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0010.001
Scholarly communication0.0010.001
Open science0.0020.002
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0030.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.027
GPT teacher head0.258
Teacher spread0.232 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)Same topicBiomedical Text Mining and OntologiesFrench-language works237,207