Bias and privacy in AI's cough-based COVID-19 recognition – Authors' reply
Bibliographic record
Abstract
We thank Humberto Perez-Espinosa and colleagues for their constructive points regarding our Comment,1Coppock H Jones L Kiskin I Schuller B COVID-19 detection from audio: seven grains of salt.Lancet Digit Health. 2021; 3: e537-38Summary Full Text Full Text PDF Scopus (21) Google Scholar which raised concerns over the work on COVID-19 detection from bioacoustic recordings. We take this opportunity to note that the study by Perez-Espinosa and colleagues2Andreu-Perez J Perez-Espinosa H Timonet E et al.A generic deep learning based cough analysis system from clinically validated samples for point-of-need COVID-19 test and severity levels.IEEE Trans Serv Comput. 2021; (published online Feb 23.)https://doi.org/10.1109/TSC.2021.3061402Crossref Scopus (42) Google Scholar represented one of the superior COVID-19 audio datasets that were collected. Although the study was not completely free from the "seven grains of salt" detailed in our Comment,1Coppock H Jones L Kiskin I Schuller B COVID-19 detection from audio: seven grains of salt.Lancet Digit Health. 2021; 3: e537-38Summary Full Text Full Text PDF Scopus (21) Google Scholar it was large scale, validated by quantitative RT-PCR, and the participants were blinded. We also applaud the recording of cycle threshold, which allowed for the comparison between model performance and viral load. We agree with the authors that participants of studies used to develop deep learning algorithms will not be able to benefit from the screening tool in a completely unbiased manner. Effort should be made to reduce this effect through careful training procedures and further scientific breakthroughs in explainable artificial intelligence and debiasing systems. In answer to the authors' concern for the privacy of participants in publicly available datasets, we admit that privacy law is not our area of expertise and we would look towards an expert in the field to comment on this. However, we note that there are a multitude of publicly available datasets containing sensitive biometric data—eg, the COVID-19 auditory respiratory dataset, COVID-19 Sounds.3Xia T Spathis D Brown C et al.COVID-19 sounds: a large-scale audio dataset for digital COVID-19 detection.https://openreview.net/forum?id=9KArJb4r5ZQDate: Aug 20, 2021Date accessed: November 2, 2021Google Scholar Nevertheless, if publication of a dataset is not possible, effort should be made to evaluate each study's model on other datasets, and to invite other research groups to evaluate their trained model on that dataset in a manner that keeps the data private and secure. We note that Perez-Espinosa and colleagues state that their "training, development (validation), and holdout (test) sets do not contain data from the same participant". This statement is a vital piece of information to include when writing up studies. The authors make an important point regarding the variability between participants; however, positive and negative cases from the same participant should still exist purely within one set and not cross train or test boundaries. We contend that absence of current published evidence does not eliminate the possibility that identity could be determined from cough audio; therefore, it cannot be considered sufficient justification for including the same individuals in both training and test sets. Given the advances of deep learning in pattern recognition, combined with audio recordings containing data besides bioacoustic information, we argue with confidence that cough recordings allow for algorithms to infer user identity to a high level. Additionally, we know from first-hand experience that when disjoint user sets are not ensured, classification performance substantially increases. This finding was shown in the study by Han and colleagues,4Han J Xia T Spathis D et al.Sounds of COVID-19: exploring realistic performance of audio-based digital testing.https://arxiv.org/abs/2106.15523Date: June 29, 2021Date accessed: November 2, 2021Google Scholar in which bias was systematically added back into the dataset. When disjoint user sets were not ensured, sensitivity scores for COVID-19 increased from 0·65 (95% CI 0·58–0·72) to 0·84 (0·75–0·92) for test users who were also present in the training set. The complexity of deep neural networks allows for memorisation of data to a degree, which results in inflated scores when not training and testing on disjoint sets.5Carlini N Liu C Erlingsson Ú Kos J Song D The secret sharer: evaluating and testing unintended memorization in neural networks.https://www.usenix.org/conference/usenixsecurity19/presentation/carliniDate accessed: November 2, 2021Google Scholar Thus, creating a disjoint test set is a fundamental prerequisite for reporting representative performance figures. Furthermore, although we agree with the authors that cough analysis is a distinct form of audio biometrics, it is within the same field of human respiratory sounds, meaning it is also subject to the issues detailed in our Comment.1Coppock H Jones L Kiskin I Schuller B COVID-19 detection from audio: seven grains of salt.Lancet Digit Health. 2021; 3: e537-38Summary Full Text Full Text PDF Scopus (21) Google Scholar We declare no competing interests. COVID-19 detection from audio: seven grains of saltDigital mass testing for COVID-19 via a mobile phone application could be made possible through machine learning and its ability to identify patterns in data. COVID-19 appears to confer unique features in the audio produced by infected individuals,1 and machine learning COVID-19 detection from breath, cough, and speech audio recordings has yielded promising results.2–4 In this critique, we present seven major issues with this research and argue that further investigation is needed before conclusions about the detectability of COVID-19 from audio can be made. Full-Text PDF Open AccessBias and privacy in AI's cough-based COVID-19 recognitionWe read with interest the Comment by Coppock and colleagues,1 in which the authors express their thoughtful opinion about several simultaneous works by independent research groups worldwide (eg, Massachusetts Institute of Technology, National Research Council of Canada, University of Cambridge, and Swiss Federal Institute of Technology Lausanne). One of these works was our own; a pioneering, multicentre, international study2 with a clinically validated dataset of forced coughs alongside quantitative RT-PCR from participants who physically attended a test centre. Full-Text PDF Open Access
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | no category Domain: not available · Genre: Commentary About the Canadian research system: no · About a Canadian topic: no | Not applicable | low |
| gpt | no category Domain: not available · Genre: Commentary About the Canadian research system: no · About a Canadian topic: no | Not applicable | high |
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".