An Introduction to the Human Connectome Project for Early Psychosis
Bibliographic record
Abstract
BACKGROUND: The time following a recent onset of psychosis is a critical period during which intervention may be maximally effective. Studying individuals in this period also offers an opportunity to investigate putative brain biomarkers of illness prior to the long-term effects of chronicity and medication. The Human Connectome Project for Early Psychosis (HCP-EP) was funded by the National Institutes of Mental Health (NIMH) as an extension of the original Human Connectome Project's approach to understanding the human brain and its structural and functional connections. DESIGN: The HCP-EP data were collected at 3 sites in Massachusetts (Beth Israel Deaconess Medical Center, McLean Hospital, and Massachusetts General Hospital), and one site in Indiana (Indiana University). Brigham and Women's Hospital served as the data coordination center and as an imaging site. RESULTS: The HCP-EP dataset includes high-quality clinical, cognitive, functional, neuroimaging, and blood specimen data acquired from 303 individuals between the ages of 16-35 years old with affective psychosis (n = 75), non-affective psychosis (n = 148), and healthy controls (n = 80). Participants with early psychosis were within 5 years of illness onset (mean duration = 1.9 years, standard deviation = 1.4 years). All data and novel or modified analytic tools developed as part of the study are publicly available to the research community through the NIMH Data Archive (NDA) or GitHub (https://github.com/pnlbwh). CONCLUSIONS: This paper provides an overview of the specific HCP-EP procedures, assessments, and protocols, as well as a brief characterization of the study participants to make it easier for researchers to use this rich dataset. Although we focus here on discussing and comparing affective and non-affective psychosis groups, the HCP-EP dataset also provides sufficient information for investigators to group participants differently.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".