European Autism GEnomics Registry (EAGER): Protocol for a multicentre cohort study and registry
Bibliographic record
Abstract
ABSTRACT Introduction Autism is a common neurodevelopmental condition with a complex genetic aetiology that includes contributions from monogenic and polygenic factors. Many autistic people have unmet healthcare needs that could be served by genomics-informed research and clinical trials. The primary aim of the European Autism GEnomics Registry (EAGER) is to establish a registry of participants with a diagnosis of autism or an associated rare genetic condition who have undergone whole-genome sequencing. The registry can facilitate recruitment for future clinical trials and research studies, based on genetic, clinical, and phenotypic profiles, as well as participant preferences. The secondary aim of EAGER is to investigate the association between mental and physical health characteristics and participants’ genetic profiles. Methods and analysis EAGER is a European multisite cohort study and registry and is part of the AIMS-2-TRIALS consortium. EAGER was developed with input from the AIMS-2-TRIALS Autism Representatives and representatives from the rare genetic conditions community. 1,500 participants with a diagnosis of autism or an associated rare genetic condition will be recruited at 13 sites across 8 countries. Participants will give a blood or saliva sample for whole-genome sequencing and answer a series of online questionnaires. Participants may also consent for the study to access pre-existing clinical data. Participants will be added to the EAGER registry. Data will be shared via the Autism Sharing Initiative, a new international collaboration aiming to create a federated system for autism data sharing. Ethics and dissemination EAGER has received full ethical approval from ethics committees in the UK (REC 23/SC/0022), Germany (S-375/2023), Portugal (CE-085/2023) and Spain (HCB/2023/0038, PIC-164-22). Approvals are in the process of being obtained from committees in Italy, Sweden, Ireland, and France. Findings will be disseminated via scientific publications and conferences, but also beyond to participants and the wider community (e.g., the AIMS-2-TRIALS website, stakeholder meetings, newsletters). STRENGHTS AND LIMITATIONS OF THIS STUDY Data from full genotyping through whole-genome sequencing will be combined with mental and physical health data and participant research priorities The EAGER sample (n=1,500), although relatively small for genetic analyses, will include a substantial proportion (around one third) of participants with a rare genetic condition, ensuring that heterogeneous presentations across the autism spectrum are captured The EAGER registry will improve the speed, efficiency, and impact of research studies and clinical trials across Europe with a culturally diverse cohort of re-contactable participants, and shared data through the Autism Sharing Initiative EAGER was developed with input from the AIMS-2-TRIALS Autism Representatives and representatives from the rare genetic conditions community Phenotypic data are collected only via self/informant-report questionnaires and not direct clinical assessments
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.071 | 0.078 |
| Meta-epidemiology (narrow) | 0.003 | 0.003 |
| Meta-epidemiology (broad) | 0.004 | 0.003 |
| Bibliometrics | 0.005 | 0.005 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.004 | 0.004 |
| Insufficient payload (model declined to judge) | 0.111 | 0.031 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".