CINECA: Common Infrastructure for National Cohorts in Europe, Canada, and Africa - Kick Off Report
Bibliographic record
Abstract
The CINECA consortium was formed in response to the EU call ‘Better Health and care, economic growth and sustainable health systems’ (H2020-SC1-BHC-2018-2020) with a proposal for an international flagship collaboration with Canada for human data storage, integration and sharing to enable personalised medicine approaches. CINECA proposes a federated cloud enabled infrastructure making population scale genomic and biomolecular data accessible across international borders, accelerating research, and improving the health of individuals across continents. CINECA will leverage international investment in human cohort studies from Europe, Canada, and Africa to deliver a paradigm shift of federated research and clinical applications. The CINECA consortium will create one of the largest cross-continental implementations of human genetic and phenotypic data federation and interoperability with a focus on common (complex) disease, one of the world’s most significant health burdens. The partners represent a unique combination of scientific excellence with experience of eleven diverse cohorts and scientific projects such as the European Genome-phenome Archive, CanDIG, and H3Africa. The CINECA Kick off meeting was held on January 24th-25th 2019 at the Wellcome Genome Campus Conference Centre, Hinxton UK. The key objective of the meeting was to bring together consortium members to facilitate discussion on the project’s goals and action plan. The report focuses on an overview of the Work Packages as presented to the consortium (focusing on deliverables due in the first reporting period), the cohorts included in the project, and the decisions made by the Executive Board for actions to implement year 1 of the project.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.046 | 0.046 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.006 | 0.006 |
| Science and technology studies | 0.006 | 0.002 |
| Scholarly communication | 0.013 | 0.008 |
| Open science | 0.008 | 0.019 |
| Research integrity | 0.004 | 0.005 |
| Insufficient payload (model declined to judge) | 0.027 | 0.017 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".