Protocol for a conversation-based analysis study: PREVENT-ED investigates dialogue features that may help predict dementia onset in later life
Bibliographic record
Abstract
INTRODUCTION: Decreasing the incidence of Alzheimer's disease (AD) is a global public health priority. Early detection of AD is an important requisite for the implementation of prevention strategies towards this goal. While it is plausible that patients at the early stages of AD may exhibit subtle behavioural signs of neurodegeneration, neuropsychological testing seems unable to detect these signs in preclinical AD. Recent studies indicate that spontaneous speech data, which can be collected frequently and naturally, provide good predictors for AD detection in cohorts with a clinical diagnosis. The potential of models based on such data for detecting preclinical AD remains unknown. METHODS AND ANALYSIS: The PREVENT-Elicitation of Dialogues (PREVENT-ED) study builds on the PREVENT Dementia project to investigate whether early behavioural signs of AD may be detected through dialogue interaction. Participants recruited through PREVENT, aged 40-59 at baseline, will be included in this study. We will use speech processing and machine learning methods to assess how well speech and visuospatial markers agree with neuropsychological, biomarker, clinical, lifestyle and genetic data from the PREVENT cohort. ETHICS AND DISSEMINATION: There are no expected risks or burdens to participants. The procedures are not invasive and do not raise significant ethical issues. We only approach healthy consenting adults and all participants will be informed that this is an exploratory study and therefore has no diagnostic aim. Confidentiality aspects such as data encryption and storage comply with the General Data Protection Regulation and with the requirements from sponsoring bodies and ethical committees. This study has been granted ethical approval by the London-Surrey Research Ethics Committee (REC reference No: 18/LO/0860), and by Caldicott and Information Governance (reference No: CRD18048). PREVENT-ED results will be published in peer-reviewed journals.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.042 | 0.057 |
| Meta-epidemiology (narrow) | 0.005 | 0.004 |
| Meta-epidemiology (broad) | 0.006 | 0.003 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.006 | 0.003 |
| Scholarly communication | 0.004 | 0.005 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.009 | 0.011 |
| Insufficient payload (model declined to judge) | 0.133 | 0.052 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".