Risk Prediction in Sexual Health Contexts: Protocol
Bibliographic record
Abstract
BACKGROUND: In British Columbia (BC), we are developing Get Checked Online (GCO), an Internet-based testing program that provides Web-based access to sexually transmitted infections (STI) testing. Much is still unknown about how to implement risk assessment and recommend tests in Web-based settings. Prediction tools have been shown to successfully increase efficiency and cost-effectiveness of STI case finding in the following settings. OBJECTIVE: This project was designed with three main objectives: (1) to derive a risk prediction rule for screening chlamydia and gonorrhea among clients attending two public sexual health clinics between 2000 and 2006 in Vancouver, BC, (2) to assess the temporal generalizability of the prediction rule among more recent visits in the Vancouver clinics (2007-2012), and (3) to assess the geographical generalizability of the rule in seven additional clinics in BC. METHODS: This study is a population-based, cross-sectional analysis of electronic records of visits collected at nine publicly funded STI clinics in BC between 2000 and 2012. We will derive a risk score from the multivariate logistic regression of clinic visit data between 2000 and 2006 at two clinics in Vancouver using newly diagnosed chlamydia and gonorrhea infections as the outcome. The area under the receiver operating characteristic curve (AUC) and the Hosmer-Lemeshow statistic will examine the model's discrimination and calibration, respectively. We will also examine the sensitivity and proportion of patients that would need to be screened at different cutoffs of the risk score. Temporal and geographical validation will be assessed using patient visit data from more recent visits (2007-2012) at the Vancouver clinics and at clinics in the rest of BC, respectively. Statistical analyses will be performed using SAS, version 9.3. RESULTS: This is an ongoing research project with initial results expected in 2014. CONCLUSIONS: The results from this research will have important implications for scaling up of Internet-based testing in BC. If a prediction rule with good calibration, discrimination, and high sensitivity to detect infection is found during this project, the prediction rule could be programmed into GCO so that the program offers individualized testing recommendations to clients. Further, the prediction rule could be adapted into educational materials to inform other Web-based content by creating awareness about STI risk factors, which may stimulate health care seeking behavior among individuals accessing the website.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.054 | 0.064 |
| Meta-epidemiology (narrow) | 0.002 | 0.003 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.005 | 0.002 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.004 | 0.008 |
| Insufficient payload (model declined to judge) | 0.164 | 0.033 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".