MétaCan
Menu
Retour à la cohorte
Enregistrement W134362049

Modeling Language Acquisition at Multiple Temporal Scales

2000· article· en· W134362049 sur OpenAlexaffabout
Steve R. Howell, Suzanna Becker

Notice bibliographique

RevueeScholarship (California Digital Library) · 2000
Typearticle
Langueen
DomaineComputer Science
ThématiqueTopic Modeling
Établissements canadiensMcMaster University
Organismes subventionnairesnon disponible
Mots-clésVerbContext (archaeology)Representation (politics)LinguisticsComputer scienceNounArgument (complex analysis)Artificial intelligenceNatural language processingCognitive sciencePsychologyCognitive psychologyHistory
DOInon disponible

Résumé

récupéré en direct d'OpenAlex

Modelling Language Acquisition at Multiple Temporal Scales Steve R. Howell (showell@hypatia.psychology.mcmaster.ca) Department of Psychology, McMaster University, 1280 Main Street West, Hamilton, Ontario, Canada Suzanna Becker (becker@mcmaster.ca) Department of Psychology, McMaster University, 1280 Main Street West, Hamilton, Ontario, Canada The problem of incorporating time in a neural network is an important one. Networks with feedback, such as Simple Recurrent Networks (SRNs) (Elman, 1990) have been argued to represent time more realistically through its effects on the processing of input, compared to standard feedforward networks. In effect, SRNs’ context units act as a memory, which incorporates a “smeared-out” representation of the network’s internal states over time. A problem does exist with this representation, however, especially for complex domains like that of language. The nature of argument agreement, embeddings and similar phenomena means that the SRN must be able to represent important past states (such as head noun for verb agreement) in spite of the declining effects of past context. While most word inputs will be related most strongly to words co- occurring close by in the input stream, verb agreement, for example, is largely determined by its corresponding noun, even in long, multiply-embedded sentences. SRN models should therefore be able to preserve representations of vital early structure for later use, in spite of the generally appropriate decline of short-term context. This issue has been addressed in an architectural fashion by others (Weckerly and Elman, 1992), but can perhaps be addressed more generally by allowing for more than one duration of context in an SRN’s operation. It is possible to apply the concept of hysteresis to the SRN’s context units. That is, the update from the hidden units to the context units may be other than the usual 1-to-1 copying; the context units may also incorporate self- recurrent connections of varying strengths. In particular, we have been experimenting with SRN’s using the hysteresis function suggested by Wermter, Arevian, and Panchev (1999) on the self-recurrent connections: Context i (t+1) = (1-Hy)*Hidden i (t) + Hy*Context i (t) We have conducted initial experiments using a test corpus derived from the original simplified test corpus used by Elman (1990). Our version differs from the original in that it includes not only consonant to vowel relations, but also word-to-word relations. That is, some of the consonant- vowel combinations (words) can only occur immediately following others. Thus in addition to the network needing to learn, for example, that u’s only come after G’s or U’s (Guuu), it must also learn that Guuu only comes after Da. It is in this capacity that the hysteresis parameter should most come into play, for it specifies, in effect, the duration of retention of the states of the context units. For short term letter to letter relations, small to zero hysteresis values should be adequate, as demonstrated originally by Elman’s success. In that experiment, the network error declined consistently within a word, but jumped at word boundaries, representing the fact that word distribution was random in that corpus. In our experiments, manipulating the hysteresis parameters was expected to bias the network in favour of either short or long term relationships. Also, simulated annealing of the learning rate, another technique not typically used with SRNs, is used in both control and experimental networks. In pilot work this feature smoothed oscillations in the gradient descent of error. The initial results of a number of simulation runs from different random initial conditions indicate that small hysteresis values (of other than 0) are indeed an advantage in learning this prediction task, with error per epoch declining noticeably, though not exceptionally, faster with 0.2 > Hys >0.1. Presumably this modest net gain is actually composed of both a larger gain for word-to-word relationships and a small decline for letter-to-letter prediction. Explorations of the exact nature of this advantage are underway, as is investigation of the best range of hysteresis parameters for various language tasks. With the ability to change the hysteresis of context layers, it becomes useful to incorporate multiple hidden layers into an SRN (Wermter, 1999), with layers having a different ‘span’ of context via different hysteresis settings. We also describe a model with multiple hidden layers that is being applied to more complex language corpora, and is designed to be able to learn at multiple time scales simultaneously, by capturing longer-range temporal structure in progressively higher layers. References Elman, J.L. (1990). Finding structure in time. Cognitive Science, 14, 179-211. Weckerly, J., & Elman, J.L. (1992). A PDP approach to processing center-embedded sentences. In Proceedings of the Fourteenth Annual Conference of the Cognitive Science Society. Hillsdale, NJ: Erlbaum. Wermter, S., Arevian, G. & Panchev, C. (1999). Recurrent Neural Network Learning for Text Routing, Proceedings of the Ninth International Conference on Artificial Neural Networks, 2, 470-475.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,000
score de la tête « metaresearch » (Gemma)0,000
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesMéta-épidémiologie (sens strict), Communication savante, Charge utile insuffisante (le modèle a refusé de juger)
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Simulation ou modélisation · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,925
Score d'incertitude au seuil1,000

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0000,000
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0020,007
Science ouverte0,0010,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0010,004

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,014
Tête enseignante GPT0,212
Écart entre enseignants0,198 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Devis d'étudeSimulation ou modélisation
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations3
Publié2000
Routes d'admission2
Résumé présentoui

Explorer davantage

Même revueeScholarship (California Digital Library)Même sujetTopic ModelingTravaux en français237 207