MétaCan
Menu
Retour à la cohorte
Enregistrement W6987686194

Towards Private Biometric Authentication and Identification

2023· dissertation· en· W6987686194 sur OpenAlexfundno aff

Notice bibliographique

RevueUWSpace (University of Waterloo) · 2023
Typedissertation
Langueen
DomaineComputer Science
ThématiqueHandwritten Text Recognition Techniques
Établissements canadiensnon disponible
Organismes subventionnairesNatural Sciences and Engineering Research Council of CanadaStrong
Mots-clésHandwritingBiometricsRobustness (evolution)Task (project management)Handwriting recognitionIdentification (biology)Set (abstract data type)Authentication (law)
DOInon disponible

Résumé

récupéré en direct d'OpenAlex

Handwriting and speech are important parts of our everyday lives. Handwriting recognition is the task that allows the recognizing of written text, whether it be letters, words or equations, from given data. When analyzing handwriting, we can analyze static images or the recording of written text through sensors. Handwriting recognition algorithms can be used in many applications, including signature verification, electronic document processing, as well as e-security and e-health related tasks. \n \nThe OnHW datasets consists of a set of datasets which, through the use of various sensors, captures the writing of characters, words, symbols and equations, recorded in the form of multivariate time series. We begin by developing character recognition models, targeting letters (and later symbols), trained and tested using the OnHW-chars dataset (and later the split OnHW-equations dataset). Our models were able to improve upon the accuracy of the previous best results on both datasets explored. Using our machine learning (ML) models, we provide 11.3%-23.56% improvements over the previous best ML models. Using deep learning (DL), as well as ensemble techniques, we were able to improve on the best previous models by 3.08%-7.01%. In addition to the accuracy improvements, we aim to provide some level of explainability, using a specialized version of LIME for time series data. This explanation helps provide some rationale for why the models make sense for the data, as well as why ensemble methods may be useful to improve accuracy rates for this task. To verify the robustness of our models trained over the OnHW-chars dataset, we trained our DL models using the same model parameters over a more recently published OnHW-equations dataset. Our DL models with ensemble learning provide 0.05%-4.75% improvements over the previous best DL models. \n \nWhile the character recognition task has many applications, when using it to provide a service, it is important to consider user privacy since handwriting is biometric data and contains private information. Next, we design a framework that uses multiparty computation (MPC) to provide users with privacy over their handwritten data, when providing a service for character recognition. We then implement the framework using the models trained on public data to provide private inference on hidden user data. This framework is implemented in the CrypTen MPC framework. We obtain results on the accuracy difference of the models when making inference using MPC, as well as the costs associated with performing this inference. We found a 0.55%-1.42% accuracy difference between plaintext inference and inference with MPC. \n \nNext, we pivot to explore writer identification, which involves identifying the writer of some handwritten text. We use the OnHW-equations dataset for our analysis, which at the time of writing has not been used for this task before. We first analyze and reformat the data to fit the writer identification task, as well as remove bias. Using DL models, we obtain accuracy results of up to 91.57% in identifying the writer using their handwriting. As with private inference in the character recognition task, it is important to account for user privacy when training writer identification models and making inference. We design and implement a framework for private training and inference for the writer recognition task, using the CrypTen MPC framework. Since training these models is very costly, we use simpler CNN's for private writer recognition. The chosen CNN trained privately in MPC obtained an accuracy of 77.45%. Next, we analyze the costs associated with privately training the CNN and other CNN's with altered model architectures. \n \nFinally, we switch to explore voice as a biometric in the speaker verification task. As with handwriting, a person's voice contains unique characteristics which can be used to determine the speaker. Not only can voice be analyzed similarly with handwriting, in that we can explore the speech recognition and speaker identification tasks, it comes with similar privacy risks for users. We design and implement a unique framework for private speaker verification using the MP-SPDZ MPC framework. We analyze the costs associated with training the model and making inferences, with our main goal being to determine the time it takes to make private inference. We then used these times as part of a survey conducted to determine how much people value the privacy of their biometrics and how long they were willing to wait for the increased privacy. We found that people were willing to tolerate significant time delays in order to privately authenticate themselves, when primed with the benefits of using MPC for privacy.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,006
score de la tête « metaresearch » (Gemma)0,011
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Simulation ou modélisation · Signal consensuel: aucune
GenreSignal candidat: Méthodes · Signal consensuel: Méthodes
Score de désaccord entre enseignants0,010
Score d'incertitude au seuil0,034

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0060,011
Méta-épidémiologie (sens strict)0,0020,001
Méta-épidémiologie (sens large)0,0020,001
Bibliométrie0,0020,002
Études des sciences et des technologies0,0010,002
Communication savante0,0040,011
Science ouverte0,0020,008
Intégrité de la recherche0,0050,005
Charge utile insuffisante (le modèle a refusé de juger)0,0100,022

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,015
Tête enseignante GPT0,232
Écart entre enseignants0,217 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSimulation ou modélisation
Domainenon disponible
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2023
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueUWSpace (University of Waterloo)Même sujetHandwritten Text Recognition TechniquesTravaux en français237 207