Quantifying dimensional severity of obsessive-compulsive disorder for neurobiological research
Bibliographic record
Abstract
Current research to explore genetic susceptibility factors in obsessive-compulsive disorder (OCD) has resulted in the tentative identification of a small number of genes. However, findings have not been readily replicated. It is now broadly accepted that a major limitation to this work is the heterogeneous nature of this disorder, and that an approach incorporating OCD symptom dimensions in a quantitative manner may be more successful in identifying both common as well as dimension-specific vulnerability genetic factors. As most existing genetic datasets did not collect specific dimensional severity ratings, a specific method to reliably extract dimensional ratings from the most widely used severity rating scale, the Yale-Brown Obsessive Compulsive Scale (YBOCS), for OCD is needed. This project aims to develop and validate a novel algorithm to extrapolate specific dimensional symptom severity ratings in OCD from the existing YBOCS for use in genetics and other neurobiological research. To accomplish this goal, we used a large data set comprising adult subjects from three independent sites: the Brazilian OCD Consortium, the Sunnybrook Health Sciences Centre in Toronto, Canada and the Hospital of Bellvitge, in Barcelona, Spain. A multinomial logistic regression was proposed to model and predict the quantitative phenotype [i.e., the severity of each of the five homogeneous symptom dimensions of the Dimensional YBOCS (DYBOCS)] in subjects who have only YBOCS (categorical) data. YBOCS and DYBOCS data obtained from 1183 subjects were used to build the model, which was tested with the leave-one-out cross-validation method. The model's goodness of fit, accepting a deviation of up to three points in the predicted DYBOCS score, varied from 78% (symmetry/order) to 84% (cleaning/contamination and hoarding dimensions). These results suggest that this algorithm may be a valuable tool for extracting dimensional phenotypic data for neurobiological studies in OCD.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.005 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".