Sex Classification Using In-ear Microphone
Bibliographic record
Abstract
In-ear devices have been getting more popular among the general public and professionals due to recent advancements in miniature sensors that enabled the incorporation of more sensors on these devices.This led to expanded applications, such as improved speech technology and communication in noisy environments such as concerts.Often, these in-ear devices block the user's ear canal at its opening with an earplug and are equipped with an In-Ear Microphone (IEM) placed inside the occluded ear canal to capture the resonance of bone and tissue-conducted speech.The motivation of this thesis is to investigate the potential use of the IEM in noisy environments in lieu of a conventional microphone placed in front of the mouth.This thesis focuses on comparing the performance of an automatic sex classifier using signals from an IEM and a microphone in front of the mouth.It first presents an overview of commonly used acoustic features in speech processing.It also covers feature optimization, adaptive denoising algorithms, and Support Vector Machine (SVM), a popular machine-learning model.An automatic speaker sex classifier pipeline, which evaluates the classification performance depending on the recording environmental condition and the microphone placement is presented.The classifier also experiments with more challenging processing conditions, such as speech-like noises and shorter input duration.The results show that the IEM captures enough acoustic features to achieve classification accuracy as high as 94.5 % in a noisy environment with the optimal quality for the training dataset and an adaptive denoising algorithm, compared to 57.7 % accuracy using the front microphone.The experimental results suggest that the IEM is a highly effective alternative to the conventional front microphone in noisy conditions.v Résumé Les dispositifs intra-auriculaires ont gagné en popularité auprès du grand public et des professionnels grâce aux récentes avancées en matière de capteurs miniatures qui ont permis d'incorporer davantage de capteurs sur ces dispositifs.Cela a conduit à des applications plus étendues, telles que l'amélioration de la technologie vocale et de la communication dans des environnements bruyants comme les concerts.Souvent, ces dispositifs intra-auriculaires bloquent le canal auditif de l'utilisateur à son ouverture avec un In-Ear Microphone (IEM) placé à l'intérieur du canal auditif occlus pour capturer la résonance de la parole transmise par les os et les tissus.La motivation de cette thèse est d'étudier l'utilisation potentielle de l'IEM dans les environnements bruyants à la place d'un microphone conventionnel placé devant la bouche.Cette thèse se concentre sur la comparaison des performances d'un classificateur automatique de genre utilisant les signaux d'un IEM et d'un microphone placé devant la bouche.Elle présente d'abord un aperçu des caractéristiques acoustiques couramment utilisées dans le traitement de la parole.Elle traite également de l'optimisation des caractéristiques, des algorithmes de débruitage adaptatif et de Support Vector Machine (SVM), un modèle populaire d'apprentissage automatique.Un pipeline de classification automatique du genre du locuteur, qui évalue les performances de classification en fonction des conditions environnementales d'enregistrement et du placement du microphone, est présenté.Le classificateur expérimente également des conditions de traitement plus difficiles, telles que des bruits de parole et une durée d'entrée plus courte.Les résultats montrent que l'IEM capture suffisamment de caractéristiques acoustiques pour atteindre une précision de classification de 94.5 % dans un environnement bruyant avec la qualité optimale pour l'ensemble de données d'entraînement et un algorithme de débruitage adaptatif, par rapport à une précision de 57.7 % en utilisant le microphone frontal.Les résultats expérimentaux suggèrent que l'IEM est une alternative très efficace au microphone frontal conventionnel dans des conditions bruyantes.vii
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.004 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".