Problematic Endorsement of Models Describing Sexual Response of Men and Women with a Sexual Partner
Bibliographic record
Abstract
In a recently published article [1], Prof. Beier and colleagues assessed their treatment program for undetected and unprosecuted help-seeking men with sexual interest toward pre-pubescents and pubescents.Introducing the research design, the authors state (p.529): "Therapy was assessed using nonrandomized waiting list control design (n = 53 treated group [TG]; n = 22 untreated control group [CG])."Self-reported changes in dynamic risk factors and offending behaviors during the 1-year treatment were used as outcome measures.The TG and the CG significantly differed in age and body age preference (see Table 1).In addition to the limitations of the research design, there are several inconsistencies in the data.In Tables 1 through 4, the authors present data for all participants who were included in the study.In contrast, in Table 5 (p.538), there are several inconsistencies in the data on the frequency and severity of offending behaviors (child sexual abuse).First, columns three and four only present data for n = 21 participants of the CG, which is inconsistent with the reported n = 22 participants in row one of Table 5.Second, there also appears to be one missing participant in the TG (n = 52 see first column of Table 5).Unfortunately, the authors do not provide any rationale for these dropouts.Because the overall sample size and thus the statistical power of the study is rather small, the authors should clarify these inconsistencies when reporting their data.To emphasize the importance of their treatment program, the authors highlight "official recidivism rates of 0%" (p.540) for all participants.However, one must consider that "current legal supervision (e.g., pending criminal charges, offense reports, criminal investigation procedures, sentencing for sexual offences involving children, and probation)" (p.531) were exclusion criteria for
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".