Logistic Regression With Machine Learning Sheds Light on the Problematic Sexual Behavior Phenotype
Bibliographic record
Abstract
OBJECTIVES: There has been a longstanding debate about whether the mechanisms involved in problematic sexual behavior (PSB) are similar to those observed in addictive disorders, or related to impulse control or to compulsivity. The aim of this report was to contribute to this debate by investigating the association between PSB, addictive disorders (internet addiction, compulsive buying), measures associated with the construct known as reward deficiency (RDS), and obsessive-compulsive disorder (OCD). METHODS: A Canadian university Office of the Registrar invited 68,846 eligible students and postdoctoral fellows. Of 4710 expressing interest in participating, 3359 completed online questionnaires, and 1801 completed the Mini-International Neuropsychiatric Interview. PSB was measured by combining those screening positive (score at least 6) on the Sexual Addiction Screening Test-Revised Core with those self-reporting PSB. Current mental health condition(s) and childhood trauma were measured by self-report. OCD was assessed by a combination of self-report and Mini-International Neuropsychiatric Interview data. RESULTS: Of 3341 participants, 407 (12.18%) screened positive on the Sexual Addiction Screening Test-Revised Core. On logistic regression, OCD, attention deficit, internet addiction, a family history of PSB, childhood trauma, compulsive buying, and male gender were associated with PSB. On multiple correspondence analysis, OCD appeared to cluster separately from the other measures, and the pattern of data differed by gender. CONCLUSIONS: In our sample, factors that have previously been associated with RDS and OCD are both associated with increased odds of PSB. The factors associated with RDS appear to contribute to a separate data cluster from OCD and to lie closer to PSB.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.066 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.005 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".