Validation of the European French Version of the Consensus Auditory-Perceptual Evaluation of Voice (CAPE-Vf)
Bibliographic record
Abstract
Objective This study aimed to validate the French adaptation of the Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V f ) for assessing voice disorders in France. The CAPE-V f addresses limitations of the GRBAS by providing a more sensitive, standardized approach to evaluating six vocal parameters (overall severity, roughness, breathiness, strain, pitch, and loudness) on three tasks (sustained vowels, sentence reading, and spontaneous speech). The study focused on investigating the intra- and inter-rater reliability, as well as the convergent and discriminant validity of the CAPE-V f . Methods Thirty-four dysphonic and seven euphonic native French speakers participated in the study. Thirteen speech-language pathologists from France evaluated the voice samples using both the CAPE-V f and GRBAS tools at a one-week interval. Intra- and inter-rater reliability were calculated using intraclass correlation coefficients (ICC), while convergent and discriminant validity were measured by correlating CAPE-V f with GRBAS and Voice Handicap Index (VHI) scores, respectively. Results The CAPE-V f showed good intra-rater reliability for overall severity (mean ICC: 0.89), strain (ICC: 0.83), and pitch (ICC: 0.88), while roughness, breathiness, and loudness exhibited moderate reliability. Inter-rater reliability was low for most parameters, except overall severity, which demonstrated good reliability (mean ICC: 0.77). Strong correlations were observed between CAPE-V f and GRBAS Grade (mean r : 0.84), supporting its convergent validity. Moderate correlations were found for roughness, breathiness, and strain. The CAPE-V f 's correlation with the VHI was moderate (mean r : 0.53), reflecting its discriminant validity. Conclusion The CAPE-V f is a valid and reliable tool for perceptual assessment of voice disorders in French-speaking populations, with stronger psychometric properties than the GRBAS, particularly for intra-rater reliability and overall severity. While inter-rater reliability was lower, qualitative feedback suggested that improvements to the protocol, particularly for pitch and loudness ratings, could enhance its clinical applicability. The findings support the CAPE-V f as a comprehensive tool for standardized clinical voice assessment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.031 | 0.063 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".