Bibliographic record
Abstract
In our work, we collected biological entities and their relations, and put them into these networks. The below description is in our manuscript. As soon as the manuscript is published, we will present the access information in this page. --- The molecule type consisted of genes, and the function type consisted of GO terms. The phenotypic type consisted of diseases that are MeSH terms in disease categories, and the concepts in some UMLS concept types: Congenital Abnormality, Acquired Abnormality, Finding, Pathologic Function, Disease or Syndrome, Mental or Behavioral Dysfunction, Cell or Molecular Dysfunction, Sign or Symptom, Anatomical Abnormality, and Neoplastic Process. Those UMLS concept types contained a considerable number of the MeSH terms in disease categories. 18 types of biological relations were defined with their pre-defined relation types and their original resources. The 18 types are presented in the Supplementary Material (Supplementary Table S2). From the CODA, 13 types of relations for this study were defined by combining the six original resources and the eleven relation types that were pre-defined in the CODA. The six original resources were BioGRID41, RegNetwork, TRANSFAC42, EndoNet43, KEGG, GO, and PhenoGO44, which were released until 2016. The six pre-defined types in CODA were: Undirected Link, Directed Link, Positive Increase, Positive Decrease, Negative Increase, and Negative Decrease. Moreover, their reversal types were also made. The Undirected Link means a non-directional association, and the Positive Decrease means that the activity or the amount of a receiver decreases as the activity or the amount of an actor increases19. From the UMLS (version 2016AA), four types of relations were defined for the present study. They were defined only by their original resources: MedlinePlus45, MTHMST46, NCI47, and OMIM. Although the UMLS has more than 200 resources, only the four resources provided a considerable number of relations among the nodes that we considered. The pre-defined types of relations in the four resources were too many for us to properly categorize them. Therefore, we ignored them when defining the types of relations for this study. The MEDLINE (version 2017) provides the co-occurrence frequency of two MeSH terms. It says how many times two MeSH terms have been attached to the same biomedical literature. Only one type of relations was defined for our research. Among functions or phenotypes, if two nodes had any co-occurrence frequency, we determined there is a co-occurrence relation between them. We tried to separate them according to their frequency. However, the performance, AUROC values, in the application was best when the co-occurrence relations were not separated.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.007 | 0.006 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.004 | 0.006 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.019 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".