Consensus-Based Definitions for Vocal Biomarkers: The International VOCAL Initiative
Bibliographic record
Abstract
Abstract Importance Voice-based health technologies are growing rapidly, but they lack standardized terminology, which hinders interdisciplinary collaboration, research quality, and clinical translation. Objective The objective of this work is to develop universally accepted definitions in the rapidly evolving field of vocal biomarkers, as part of the VOCAL ( V ocal Biomarker Guidelines for O ntology, C lassification, A pplication, and L ogistics) initiative, a structured, international consensus-based framework that aims to provide standards, and guidelines. Design VOCAL is a rigorous, international, multi-stage consensus-building study conducted in 2024-2025. Setting Multi-institutional collaboration between representatives from the Bridge2AI-Voice Consortium (North America) and the eVoiceNet Network (European Union), culminating in an in-person workshop at the 2025 Bridge2AI Voice Symposium. Participants A group of 24 international experts in medicine, clinical research, speech and language, audio signal processing, statistics, methodology, regulation, ethics. Methods VOCAL’s iterative process involved five rounds of review, feedback, and an in-person workshop at the international 2025 Bridge2AI Voice Symposium, ensuring the incorporation of diverse perspectives and achieving a robust agreement on the proposed definitions. Main outcomes and measures Consensus-based definitions for vocal biomarkers, spanning from broad concepts (biomarker, digital biomarker, vocal biomarker) to domain-specific measures (cardio-respiratory acoustic, voice, speech/articulatory, cognitive/language). Results A hierarchical continuum model of vocal biomarkers was established. We first distinguished between the concepts of vocal measures and vocal biomarkers. We then defined terms from broad, overarching concepts (Level 0: Biomarker, Digital Biomarker, Vocal Biomarker) to more specific physiological and cognitive domains (Level 1: Cardio-Respiratory Acoustic; Level 2: Voice; Level 3: Speech/Articulatory; Level 4: Cognitive/Language, including linguistic and paralinguistic subtypes). Conclusions and Relevance This work provides a shared vocabulary that is essential for fostering communication through interdisciplinary collaboration, improving the quality and efficiency of research and development, and ensuring the ethical, reliable, and scalable deployment of future voice-based health technologies. It lays foundational groundwork for upcoming guidelines and standards, which are crucial for advancing the field of vocal biomarkers into widespread clinical utility. Key points Question Can we establish an international consensus on the definitions related to vocal biomarkers? Findings Through a multi-stage international consensus process involving expert representatives from the Bridge2AI-Voice Consortium and eVoiceNet Network, a hierarchical model of vocal biomarkers was developed, defining terms from broad concepts (biomarker, digital biomarker, vocal biomarker) to specific physiological and cognitive domains (cardio-respiratory acoustic, voice, speech/articulatory, cognitive/language). Meaning Standardized definitions for vocal biomarkers provide an essential shared vocabulary for interdisciplinary collaboration and lay the foundational groundwork for future guidelines needed to advance the implementation of voice-based health technologies into clinical practice and clinical research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".