Gene-Centric Database Reveals Environmental and Lifestyle Relationships for Potential Risk Modification and Prevention
Bibliographic record
Abstract
The database at Nutrigenetics.net has been under development since 2007 to facilitate the identification and classification of PubMed articles relevant to human genetics. A controlled vocabulary (i.e., standardized terminology) is used to index these records, with links back to PubMed for every article title. This enables the display of indexes (alphabetical subtopic listings) for any given topic, or for any given combination of topics, including for genes and specific genetic variants. Stepwise use of such indexes (first for one topic, then for combinations of topics) can reveal relationships that are otherwise easily overlooked. These relationships include environmental and lifestyle variables with potential relevance to risk modification (both beneficial and detrimental), and to prevention, or at least to the potential delay of symptom onset for health conditions like Alzheimer disease among many others. Thirty-four specific genetic variants have each been mentioned in at least ≥1,000 PubMed titles/abstracts, and these numbers are steadily increasing. The benefits of indexing with standardized terminology are illustrated for genetic variants like MTHFR 677C-T and its various synonyms (e.g., rs1801133 or Ala222Val). Such use of a controlled vocabulary is also helpful for numerous health conditions, and for potential risk modifiers (i.e., potential risk/effect modifiers).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".