A Multi-Agent Approach to Generating Context-Rich Gene Sets
Bibliographic record
Abstract
Abstract Gene sets are collections of genes that share a common biological function, process, or component that can be used to get insight into the biological relevance of genomic data. Databases containing these gene sets aids in a wide array of analytical methods. The results of these methods, such as gene set analysis or phenotype-based gene prioritization, depend on the quality of the gene sets. Despite the extensive literature and genetic data available for constructing these databases, they often lack sufficient biological context. Current curation methods rely on labour-intensive expert manual curation from literature and datasets, as well as automated methods that are not context-aware. Therefore, there is a significant opportunity to utilize publicly available literature to bridge this gap and create more precise gene sets. With the advancement of natural language processing technologies, particularly large language models, this task can be performed more efficiently. In this work, we present a multi-agent system that utilizes the Llama 3, DeepSeek, and Qwen open-source large language models to analyze PubMed abstracts, allowing us to reconstruct gene sets in existing databases that better reflect specific biological contexts. Our approach consists of two pipelines. One verifies the inclusion of genes in a gene set by proof of evidence in the abstracts showing the association between the gene and the gene set. The second pipeline parses through the abstracts to identify genes not already included in the gene set for potential inclusion. To evaluate the proposed approach, we reconstructed a random selection of gene sets within the Human Ontology Phenotype (HPO). Our analysis shows that 149 of these gene sets have a similarity of 65.18% when compared to the original HPO gene sets, aligning well with the current HPO database. Additionally, we found an average of 3.15 new genes not included in the HPO gene sets, each supported by verified literature linking them to their respective gene sets. This highlights that our updated gene set database better reflects the current state of biological findings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".