A graph-embedded topic model enables characterization of diverse pain phenotypes among UK Biobank individuals
Bibliographic record
Abstract
SUMMARY Large biobank repositories of clinical conditions and medications data open opportunities to investigate the phenotypic disease network. To enable systematic investigation of entire structured phenomes, we present graph embedded topic model (GETM). Our contributions are two folds in terms of method and applications. On the methodology side, we offer two main contributions in GETM. First, to aid topic inference, we integrate existing biomedical knowledge graph information in the form of pre-trained graph embedding into the embedded topic model. Second, leveraging deep learning techniques, we developed a variational autoencoder framework to infer patient phenotypic mixture by modeling multi-modal discrete patient medical records. In particular, for interpretability, we use a linear decoder to simultaneously infer the bi-modal distributions of the disease conditions and medications. On the application side, we applied GETM to UK Biobank (UKB) self-reported clinical phenotype data, which contains 443 self-reported medical conditions and 802 self-reported medications for 457,461 individuals. Compared to existing methods, GETM demonstrates overall superior performance in imputing missing conditions and medications. Here, we focused on characterizing pain phenotypes recorded in the questionnaire of the UKB individuals. GETM accurately predicts the status of chronic musculoskeletal (CMK) pain, chronic pain by body-site, and non-specific chronic pain using past conditions and medications. Our analyses revealed not only the known pain-related topics but also the surprising predominance of medications and conditions in the cardiovascular category among the most predictive topics across chronic pain phenotypes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".