MétaCan
Menu
Retour à la cohorte
Enregistrement W3129585455 · doi:10.5075/epfl-thesis-10404

The Learnability of the Grammar of Jazz: Bayesian Inference of Hierarchical Structures in Harmony

2020· article· en· W3129585455 sur OpenAlexfundno aff
Daniel Harasim

Notice bibliographique

RevueInfoscience (Ecole Polytechnique Fédérale de Lausanne) · 2020
Typearticle
Langueen
DomaineComputer Science
ThématiqueMusic and Audio Processing
Établissements canadiensnon disponible
Organismes subventionnairesNatural Sciences and Engineering Research Council of CanadaTechnische Universität DresdenDeutsche ForschungsgemeinschaftEuropean Commission
Mots-clésLearnabilityHarmony (color)Artificial intelligenceInferenceGrammarNatural language processingBayesian inferenceGrammar inductionComputer scienceLinguisticsMathematicsBayesian probabilityPhilosophyArt

Résumé

récupéré en direct d'OpenAlex

Musical grammar describes a set of principles that are used to understand and interpret the structure of a piece according to a musical style. The main topic of this study is grammar induction for harmony --- the process of learning structural principles from the observation of chord sequences. The question how grammars are learnable by induction from sequential data is an instance of the more general question how abstract knowledge is inducible from the observation of data --- a central question of cognitive science. Under the assumption that human learning approximately follows the principles of rational reasoning, Bayesian models of cognition can be used to simulate learning processes. This study investigates what prior knowledge makes it possible to learn musical grammar inductively from Jazz chord sequences using Bayesian models and computational simulations. The theoretical part of the thesis presents how questions about learnability can be studied in a unified framework involving music analysis, cognitive modeling, Bayesian statistics, and computational simulations. A new grammar formalism, called Probabilistic Abstract Context-Free Grammar (PACFG), is proposed that allows for flexible probability models which facilitate the grammar-induction experiments of this study. PACFG can jointly model multiple musical dimensions such as harmony and rhythm, and can use coordinate ascent variational inference for grammar learning. The empirical part of the thesis reports supervised and unsupervised grammar-learning experiments. To train and evaluate grammar models, a ground-truth dataset of hierarchical analyses of complete Jazz standards, called the Jazz Harmony Treebank (JHT), was created. The supervised grammar-learning experiments, in which grammars for Jazz harmony are learned from the JHT analyses, show that jointly modeling harmony and rhythm significantly improves the grammar models' prediction of the ground truth. The performance and robustness of the grammars are further improved by a transpositionally invariant parameterization of rule probabilities. Following the supervised grammar learning, unsupervised grammar learning was performed by inducing harmony grammars merely from Jazz chord sequences, without the observation of the JHT trees. The results show that the best induced grammar performs similarly well as the best supervised grammar. In particular, the goal-directedness of functional harmony does not need to be assumed a priori, but can be learned without usage of music-specific prior knowledge. The findings of this thesis show that general prior knowledge enables an ideal learner to acquire abstract musical principles by statistical learning. In conclusion, it is plausible that much aspects of musical grammar have been learned by Jazz musicians and listeners, instead of being innate predispositions or explicitly taught concepts. This thesis is moreover embedded into the context of empirical music research and digital humanities. Current studies either describe complex musical structures qualitatively or investigate simpler aspects quantitatively. The computational models developed in this thesis demonstrate that deep insights into music and statistical analyses are not mutually exclusive. They enable a new kind of data-driven music theory and musicology, for instance through comparative analyses of musical grammar for different styles such as Jazz, Rock, and Western classical music.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,007
score de la tête « metaresearch » (Gemma)0,053
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Simulation ou modélisation · Signal consensuel: Simulation ou modélisation
GenreSignal candidat: Empirique · Signal consensuel: aucune
Score de désaccord entre enseignants0,007
Score d'incertitude au seuil0,037

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0070,053
Méta-épidémiologie (sens strict)0,0000,001
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0010,000
Études des sciences et des technologies0,0010,003
Communication savante0,0020,005
Science ouverte0,0010,002
Intégrité de la recherche0,0010,004
Charge utile insuffisante (le modèle a refusé de juger)0,0020,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,016
Tête enseignante GPT0,257
Écart entre enseignants0,241 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSimulation ou modélisation
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations3
Publié2020
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueInfoscience (Ecole Polytechnique Fédérale de Lausanne)Même sujetMusic and Audio ProcessingTravaux en français237 207