Development of a computable phenotype to identify a transgender sample for health research purposes: a feasibility study in a large linked provincial healthcare administrative cohort in British Columbia, Canada
Bibliographic record
Abstract
OBJECTIVES: Innovative methods are needed for identification of transgender people in administrative records for health research purposes. This study investigated the feasibility of using transgender-specific healthcare utilisation in a Canadian population-based health records database to develop a computable phenotype (CP) and identify the proportion of transgender people within the HIV-positive population as a public health priority. DESIGN: The Comparative Outcomes and Service Utilization Trends (COAST) Study cohort comprises a data linkage between two provincial data sources: The British Columbia (BC) Centre for Excellence in HIV/AIDS Drug Treatment Program, which coordinates HIV treatment dispensation across BC and Population Data BC, a provincial data repository holding individual, longitudinal data for all BC residents (1996-2013). SETTING: British Columbia, Canada. PARTICIPANTS: COAST participants include 13 907 BC residents living with HIV (≥19 years of age) and a 10% random sample comparison group of the HIV-negative general population (514 952 individuals). PRIMARY AND SECONDARY OUTCOME MEASURES: Healthcare records were used to identify transgender people via a CP algorithm (diagnosis codes+androgen blocker/hormone prescriptions), to examine related diagnoses and prescription concordance and to validate the CP using an independent provider-reported transgender status measure. Demographics and chronic illness burden were also characterised for the transgender sample. RESULTS: The best-performing CP identified 137 HIV-negative and 51 HIV-positive transgender people (total 188). In validity analyses, the best-performing CP had low sensitivity (27.5%, 95% CI: 17.8% to 39.8%), high specificity (99.8%, 95% CI: 99.6% to 99.8%), low agreement using Kappa statistics (0.3, 95% CI: 0.2 to 0.5) and moderate positive predictive value (43.2%, 95% CI: 28.7% to 58.9%). There was high concordance between exogenous sex hormone use and transgender-specific diagnoses. CONCLUSIONS: The development of a validated CP opens up new opportunities for identifying transgender people for inclusion in population-based health research using administrative health data, and offers the potential for much-needed and heretofore unavailable evidence on health status, including HIV status, and the healthcare use and needs of transgender people.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.017 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.005 | 0.001 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".