MétaCan
Menu
Retour à la cohorte
Enregistrement W3215872957 · doi:10.1016/s2589-7500(21)00233-8

Bias and privacy in AI's cough-based COVID-19 recognition – Authors' reply

2021· letter· en· W3215872957 sur OpenAlexaboutno aff
Harry Coppock, Lyn Jones, Ivan Kiskin, Björn W. Schuller

Notice bibliographique

RevueThe Lancet Digital Health · 2021
Typeletter
Langueen
DomaineMedicine
ThématiqueCOVID-19 diagnosis using AI
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésCoronavirus disease 2019 (COVID-19)Scopus2019-20 coronavirus outbreakSevere acute respiratory syndrome coronavirus 2 (SARS-CoV-2)Artificial intelligenceTest (biology)Computer scienceMedicineMachine learningMEDLINEInternal medicinePolitical sciencePathologyBiologyLaw

Résumé

récupéré en direct d'OpenAlex

We thank Humberto Perez-Espinosa and colleagues for their constructive points regarding our Comment,1Coppock H Jones L Kiskin I Schuller B COVID-19 detection from audio: seven grains of salt.Lancet Digit Health. 2021; 3: e537-38Summary Full Text Full Text PDF Scopus (21) Google Scholar which raised concerns over the work on COVID-19 detection from bioacoustic recordings. We take this opportunity to note that the study by Perez-Espinosa and colleagues2Andreu-Perez J Perez-Espinosa H Timonet E et al.A generic deep learning based cough analysis system from clinically validated samples for point-of-need COVID-19 test and severity levels.IEEE Trans Serv Comput. 2021; (published online Feb 23.)https://doi.org/10.1109/TSC.2021.3061402Crossref Scopus (42) Google Scholar represented one of the superior COVID-19 audio datasets that were collected. Although the study was not completely free from the "seven grains of salt" detailed in our Comment,1Coppock H Jones L Kiskin I Schuller B COVID-19 detection from audio: seven grains of salt.Lancet Digit Health. 2021; 3: e537-38Summary Full Text Full Text PDF Scopus (21) Google Scholar it was large scale, validated by quantitative RT-PCR, and the participants were blinded. We also applaud the recording of cycle threshold, which allowed for the comparison between model performance and viral load. We agree with the authors that participants of studies used to develop deep learning algorithms will not be able to benefit from the screening tool in a completely unbiased manner. Effort should be made to reduce this effect through careful training procedures and further scientific breakthroughs in explainable artificial intelligence and debiasing systems. In answer to the authors' concern for the privacy of participants in publicly available datasets, we admit that privacy law is not our area of expertise and we would look towards an expert in the field to comment on this. However, we note that there are a multitude of publicly available datasets containing sensitive biometric data—eg, the COVID-19 auditory respiratory dataset, COVID-19 Sounds.3Xia T Spathis D Brown C et al.COVID-19 sounds: a large-scale audio dataset for digital COVID-19 detection.https://openreview.net/forum?id=9KArJb4r5ZQDate: Aug 20, 2021Date accessed: November 2, 2021Google Scholar Nevertheless, if publication of a dataset is not possible, effort should be made to evaluate each study's model on other datasets, and to invite other research groups to evaluate their trained model on that dataset in a manner that keeps the data private and secure. We note that Perez-Espinosa and colleagues state that their "training, development (validation), and holdout (test) sets do not contain data from the same participant". This statement is a vital piece of information to include when writing up studies. The authors make an important point regarding the variability between participants; however, positive and negative cases from the same participant should still exist purely within one set and not cross train or test boundaries. We contend that absence of current published evidence does not eliminate the possibility that identity could be determined from cough audio; therefore, it cannot be considered sufficient justification for including the same individuals in both training and test sets. Given the advances of deep learning in pattern recognition, combined with audio recordings containing data besides bioacoustic information, we argue with confidence that cough recordings allow for algorithms to infer user identity to a high level. Additionally, we know from first-hand experience that when disjoint user sets are not ensured, classification performance substantially increases. This finding was shown in the study by Han and colleagues,4Han J Xia T Spathis D et al.Sounds of COVID-19: exploring realistic performance of audio-based digital testing.https://arxiv.org/abs/2106.15523Date: June 29, 2021Date accessed: November 2, 2021Google Scholar in which bias was systematically added back into the dataset. When disjoint user sets were not ensured, sensitivity scores for COVID-19 increased from 0·65 (95% CI 0·58–0·72) to 0·84 (0·75–0·92) for test users who were also present in the training set. The complexity of deep neural networks allows for memorisation of data to a degree, which results in inflated scores when not training and testing on disjoint sets.5Carlini N Liu C Erlingsson Ú Kos J Song D The secret sharer: evaluating and testing unintended memorization in neural networks.https://www.usenix.org/conference/usenixsecurity19/presentation/carliniDate accessed: November 2, 2021Google Scholar Thus, creating a disjoint test set is a fundamental prerequisite for reporting representative performance figures. Furthermore, although we agree with the authors that cough analysis is a distinct form of audio biometrics, it is within the same field of human respiratory sounds, meaning it is also subject to the issues detailed in our Comment.1Coppock H Jones L Kiskin I Schuller B COVID-19 detection from audio: seven grains of salt.Lancet Digit Health. 2021; 3: e537-38Summary Full Text Full Text PDF Scopus (21) Google Scholar We declare no competing interests. COVID-19 detection from audio: seven grains of saltDigital mass testing for COVID-19 via a mobile phone application could be made possible through machine learning and its ability to identify patterns in data. COVID-19 appears to confer unique features in the audio produced by infected individuals,1 and machine learning COVID-19 detection from breath, cough, and speech audio recordings has yielded promising results.2–4 In this critique, we present seven major issues with this research and argue that further investigation is needed before conclusions about the detectability of COVID-19 from audio can be made. Full-Text PDF Open AccessBias and privacy in AI's cough-based COVID-19 recognitionWe read with interest the Comment by Coppock and colleagues,1 in which the authors express their thoughtful opinion about several simultaneous works by independent research groups worldwide (eg, Massachusetts Institute of Technology, National Research Council of Canada, University of Cambridge, and Swiss Federal Institute of Technology Lausanne). One of these works was our own; a pioneering, multicentre, international study2 with a clinically validated dataset of forced coughs alongside quantitative RT-PCR from participants who physically attended a test centre. Full-Text PDF Open Access

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Étiquettes directes de modèles (non validées)

Étiquettes de catégorie et de devis d'étude par modèle, issues des rondes d'étiquetage. C'est une sortie machine, non validée, et le désaccord entre modèles est livré comme donnée. Aucun devis ici n'est encore validé contre MEDLINE.

BrasCatégoriesDevis d'étudeConfiance
gemmaaucune catégorie
Domaine: non disponible · Genre: Commentaire
Porte sur le système de recherche canadien: non · Porte sur un sujet canadien: non
Sans objetlow
gptaucune catégorie
Domaine: non disponible · Genre: Commentaire
Porte sur le système de recherche canadien: non · Porte sur un sujet canadien: non
Sans objethigh
modèles en accordL'accord compare des ensembles de catégories et des devis identiques entre les bras.

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,001
score de la tête « metaresearch » (Gemma)0,004
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesMéta-épidémiologie (sens strict), Intégrité de la recherche
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Commentaire · Signal consensuel: Commentaire
Score de désaccord entre enseignants0,010
Score d'incertitude au seuil1,000

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0010,004
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0010,000
Bibliométrie0,0000,001
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,003
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,178
Tête enseignante GPT0,399
Écart entre enseignants0,222 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Étiqueté directement par 2 modèles lisant le dossier complet.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreCommentaire

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations1
Publié2021
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueThe Lancet Digital HealthMême sujetCOVID-19 diagnosis using AITravaux en français237 207