P1-S4.06 What impact does missing Quebec data have on national HIV surveillance data?
Bibliographic record
Abstract
Objective To quantify the difference in the exposure category breakdowns of national HIV surveillance figures if exposure data from the Institut nationale de Santé Publique du Québec (INSPQ) were included in national datasets. Background National HIV/AIDS surveillance is coordinated by the Public Health Agency of Canada's (PHAC) Surveillance and Risk Assessment Division's (SRAD). HIV is reportable in all provinces and territories, although the degree of epidemiologic information collected and submitted varies. Quebec's case reports to PHAC come from their laboratory-based surveillance system, which contains positive test reports, by age and sex. All Quebec cases are classified in SRAD's dataset as Not Reported, which contributes to the large proportion of cases at the national level with no known exposure category. Methods Quebec's provincial HIV surveillance system “Programme de surveillance de l'infection par le VIH au Québec” collects further epidemiological information, including exposure category and risk factor information, although recorded separately from the HIV laboratory test results file. This provincial system's exposure category data was added to existing national surveillance data, and the exposure category breakdowns recalculated, in order to assess change in the proportion of unknown/not reported cases and to quantify the resulting difference in exposure category breakdowns at the national level. Results With inclusion of Quebec data for 2009, there is a 50% decrease (from 45.5% to 23.1%) in the proportion of national HIV cases with unknown exposure category. There are also differences in the overall national exposure category breakdowns. For 2009, proportional increases were observed in the men who have sex with men (MSM) and heterosexual-endemic categories (5.4% and 2.8% respectively), while proportional decreases were observed in the exposure categories of injection drug use (−4.1%), heterosexual-risk (−2.0%), and no-identified-risk heterosexual (−2.2%). Conclusions Inclusion of Quebec's risk exposure data in the national HIV dataset is significant; the national dataset becomes more complete and the proportion of cases with unknown exposure category is reduced. This analysis demonstrates that inclusion of exposure category data, from the provincial HIV surveillance system of Quebec's INSPQ can alter the exposure category breakdowns at the national level, thereby offering a more accurate picture of HIV diagnoses in Canada.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.002 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".