14. IS BIGGER BETTER? PROMISES AND PITFALLS OF BIG DATA IN NEUROIMAGING OF PSYCHOSIS
Bibliographic record
Abstract
Big neuroimaging datasets comprising hundreds or even thousands of subjects are becoming widely available, thanks to major collaborative efforts across multiple imaging centers and groups. Mining and analyzing Big-Data is also becoming feasible, owing to increased computational power and new implementations of machine learning algorithms, which can learn from data and generate predictions. Big-Data studies bear exceptional promise in disentangling complex psychiatric illness, including psychosis, where imaging correlates are often subtle and difficult to reproduce. Large datasets, combined with novel machine learning algorithms have opened avenues for delineating subtypes, as well as for predicting biological and clinical outcomes, including psychosis conversion and drug response. However, the use of Big-Data generates challenges in terms of design, construction and statistical analyses. One such problem presented by current studies relates to harmonizing neuroimaging signals across multiple centers or sites, as no current gold standard exists. Moreover, harmonizing disparate populations could introduce unwanted covariables, and potentially attenuate psychosis-related effects. Additional challenge relates to incomplete or incompatible clinical information between different study sites, which diminishes usable clinical measures and in turn, the clinical scope of Big-data studies. This symposium is designed to address the practical and theoretical aspects of acquiring and analyzing large neuroimaging studies in psychosis and schizophrenia-spectrum disorders. We will describe new and exciting opportunities that come with Big-Data studies, as well as the pitfalls limiting current methods. Throughout, we will provide practical recommendations to maximize the potential of big-data. In addition, we will explore the current tradeoff between Big-data studies that bolster sensitivity and smaller studies, which enable nuanced investigation of homogenous populations and specific imaging markers to elucidate underlying pathologies in psychosis. We argue that both Big and small data play important roles in our effort to understand psychosis and its treatment. This symposium will bring together five leading neuroimagers, who will present different aspects of Big versus small data studies, while providing critical cautionary remarks regarding the shortcomings of these methods: 1) Prof. Neda Jahanshad, PhD, of the Keck School of Medicine, University of Southern California, is a key player in the ENIGMA network, and its Big-Data that was constructed through meta analyses. She will present key findings from the ENIGMA studies, as well as important technical considerations that arise in the meta-analysis process, including combining clinical information across multiple study sites. 2) Prof. Bo (Cloud) Cao, PhD, of the Department of Psychiatry, University of Alberta, Canada, is an expert in the modification and application of machine learning approaches for imaging studies. He will present state-of-the-art prediction algorithms and describe the advantages of combining these algorithms with Big-Data in psychosis studies. 3) Prof. Jennifer Coughlin, MD, of the Department of Psychiatry and Behavioral Sciences, Johns Hopkins University applies novel PET radioligands that are developed to specifically target molecules relevant to biological pathophysiology. She will present her work in psychosis, highlighting the important role of smaller, yet more specific studies. 4) Prof. Ofer Pasternak, PhD, of the Departments of Psychiatry and Radiology, Harvard Medical School, is a developer of more-specific MRI measures, and has been designing acquisition protocols for large multi-site studies. He will present recent advances in the harmonization of diffusion MRI data, and discuss the trade-off between small imaging studies and large harmonized imaging studies. The discussant will be Prof. Carrie Bearden, PhD. Department of Psychology, UCLA. Dr. Bearden had leading roles in a number of large imaging studies (e.g., NAPLS, ENIGMA, ABCD) as well as smaller studies. She will complement the panel by bringing in a more clinical point of view, informed of the practical needs of neuroscientists who are considering entering large multi-site studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.125 | 0.210 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.004 | 0.018 |
| Scholarly communication | 0.016 | 0.053 |
| Open science | 0.005 | 0.011 |
| Research integrity | 0.012 | 0.019 |
| Insufficient payload (model declined to judge) | 0.013 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".