Multivariate consistency of resting-state fMRI connectivity maps acquired on a single individual over 2.5 years, 13 sites and 3 vendors
Bibliographic record
Abstract
Studies using resting-state functional magnetic resonance imaging (rsfMRI) are increasingly collecting data at multiple sites in order to speed up recruitment or increase sample size. Multisite studies potentially introduce systematic biases in connectivity measures across sites, which may negatively impact the detection of clinical effects. Long-term multisite biases (i.e. over several years) are still poorly understood. The main objective of this study was to assess the long-term consistency of rsfMRI multisite connectivity measures derived from the harmonized Canadian Dementia Imaging Protocol (CDIP, www.cdip-pcid.ca). Nine to ten minutes of functional BOLD images were acquired from an adult cognitively healthy volunteer scanned repeatedly at 13 Canadian sites on three scanner makes (General Electric, Philips and Siemens) over the course of 2.5 years. RsfMRI connectivity maps were extracted for each session in seven canonical functional networks. The reliability (spatial Pearson’s correlation) of maps was about 0.6, with moderate effects (up to 0.2) of scanner makes and sites. The time elapsed between scans had a negligible effect on the consistency of connectivity maps. To assess the utility of such measures in machine learning models, we pooled the long-term longitudinal data with a single-site, short-term (1 month) data sample acquired on 26 subjects (10 scans per subject), called HNU1. Using randomly selected pairs of scans from each subject, we quantified the ability of a data-driven unsupervised cluster analysis to match the two scans. In this “fingerprinting” experiment, we found that scans from the Canadian subject could be matched with high accuracy (>85% for some networks), and fell in the range of accuracies observed for the HNU1 subjects. Overall, these results support the feasibility of multivariate, machine learning analysis of rsfMRI measures in a multisite study that extends for several years, even with fairly short (approximately ten minutes) time series.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".