Methodology for Biological Sample Collection, Processing, and Storage in the Newcastle 1000 Pregnancy Cohort: Protocol for a Longitudinal, Prospective Population-Based Study in Australia
Bibliographic record
Abstract
BACKGROUND: Research in the developmental origins of health and disease provides compelling evidence that adverse events during the first 1000 days of life from conception can impact life course health. Despite many decades of research, we still lack a complete understanding of the mechanisms underlying some of these associations. The Newcastle 1000 Study (NEW1000) is a comprehensive, prospective population-based pregnancy cohort study based in Newcastle, New South Wales, Australia, that will recruit pregnant women and their partners at 11-14 weeks' gestation, with assessments at 20, 28, and 36 weeks; birth; 6 weeks; and 6 months, in order to provide detailed data about the first 1000 days of life to investigate the developmental origins of noncommunicable diseases. OBJECTIVE: The study aims to provide a longitudinal multisystem approach to phenotyping, supported by robust clinical data and collection of biological samples in NEW1000. METHODS: This manuscript describes in detail the large variety of samples collected in the study and the method of collection, storage, and utility of the samples in the biobank, with a particular focus on incorporation of the samples into emerging and novel large-scale "-omics" platforms, including the genome, microbiome, epigenome, transcriptome, fragmentome, metabolome, proteome, exposome, and cell-free DNA and RNA. Specifically, this manuscript details the methods used to collect, process, and store biological samples, including maternal, paternal, and fetal blood, microbiome (stool, skin, vaginal, oral), urine, saliva, hair, toenail, placenta, colostrum, and breastmilk. RESULTS: Recruitment for the study began in March 2021. As of July 2024, 1040 women and 684 partners were enrolled, with 922 infants born. The NEW1000 biobank contains 24,357 plasma aliquots from ethylenediaminetetraacetic acid (EDTA) tubes, 5284 buffy coat aliquots, 4000 plasma aliquots from lithium heparin tubes, 15,884 blood serum aliquots, 2977 PAX RNA tubes, 26,595 urine sample aliquots, 2280 fecal swabs, 17,687 microbiome swabs, 2356 saliva sample aliquots, 1195 breastmilk sample aliquots, 4007 placental tissue aliquots, 2680 hair samples, and 2193 nail samples. CONCLUSIONS: NEW1000 will generate a multigenerational, deeply phenotyped cohort with a comprehensive biobank of samples relevant to a large variety of analyses, including multiple -omics platforms. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/63562.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.042 | 0.046 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.004 | 0.002 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.045 | 0.014 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".