Population Size Estimation of Gay and Bisexual Men and Other Men Who Have Sex With Men Using Social Media-Based Platforms
Bibliographic record
Abstract
BACKGROUND: Gay, bisexual, and other cisgender men who have sex with men (GBMSM) are disproportionately affected by the HIV pandemic. Traditionally, GBMSM have been deemed less relevant in HIV epidemics in low- and middle-income settings where HIV epidemics are more generalized. This is due (in part) to how important population size estimates regarding the number of individuals who identify as GBMSM are to informing the development and monitoring of HIV prevention, treatment, and care programs and coverage. However, pervasive stigma and criminalization of same-sex practices and relationships provide a challenging environment for population enumeration, and these factors have been associated with implausibly low or absent size estimates of GBMSM, thereby limiting knowledge about the dynamics of HIV transmission and the implementation of programs addressing GBMSM. OBJECTIVE: This study leverages estimates of the number of members of a social app geared towards gay men (Hornet) and members of Facebook using self-reported relationship interests in men, men and women, and those with at least one reported same-sex interest. Results were categorized by country of residence to validate official size estimates of GBMSM in 13 countries across five continents. METHODS: Data were collected through the Hornet Gay Social Network and by using an a priori determined framework to estimate the numbers of Facebook members with interests associated with GBMSM in South Africa, Ghana, Nigeria, Senegal, Côte d'Ivoire, Mauritania, The Gambia, Lebanon, Thailand, Malaysia, Brazil, Ukraine, and the United States. These estimates were compared with the most recent Joint United Nations Programme on HIV/AIDS (UNAIDS) and national estimates across 143 countries. RESULTS: The estimates that leveraged social media apps for the number of GBMSM across countries are consistently far higher than official UNAIDS estimates. Using Facebook, it is also feasible to assess the numbers of GBMSM aged 13-17 years, which demonstrate similar proportions to those of older men. There is greater consistency in Facebook estimates of GBMSM compared to UNAIDS-reported estimates across countries. CONCLUSIONS: The ability to use social media for epidemiologic and HIV prevention, treatment, and care needs continues to improve. Here, a method leveraging different categories of same-sex interests on Facebook, combined with a specific gay-oriented app (Hornet), demonstrated significantly higher estimates than those officially reported. While there are biases in this approach, these data reinforce the need for multiple methods to be used to count the number of GBMSM (especially in more stigmatizing settings) to better inform mathematical models and the scale of HIV program coverage. Moreover, these estimates can inform programs for those aged 13-17 years; a group for which HIV incidence is the highest and HIV prevention program coverage, including the availability of pre-exposure prophylaxis (PrEP), is lowest. Taken together, these results highlight the potential for social media to provide comparable estimates of the number of GBMSM across a large range of countries, including some with no reported estimates.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".