Notice bibliographique
Résumé
G eneralization In Sim ple R ecurrent N etw orks M arius V ilcu (m vilcu@ cs.sfu.ca) School of Com puting Science Sim on Fraser U niversity, 8888 U niversity D rive, Burnaby, Canada, V 5A 1S6 R obert F. H adley (hadley@ cs.sfu.ca) School of Com puting Science Sim on Fraser U niversity, 8888 U niversity D rive, Burnaby, Canada, V 5A 1S6 A bstract In this paper w e exam ine Elm an’s position (1999) on generalization in sim ple recurrent netw orks. Elm an’s sim ulation is a response to M arcus et al.’s (1999) experim ent w ith infants; specifically their ability to differentiate betw een novel sequences of syllables of the form A BA and A BB. Elm an contends that SRN s can learn to generalize to novel stim uli, just as M arcus et al’s infants did. H ow ever, w e believe that Elm an’s conclusions are overstated. Specifically, w e perform ed large batch experim ents involving sim ple recurrent netw orks w ith differing data sets. O ur results show ed that SRN s are m uch less successful than Elm an asserted, although there is a w eak tendency for netw orks to respond m eaningfully, rather than random ly, to input stim uli. Introduction In a recent paper, Elm an (1999) casts doubt upon the w idely noted results of M arcus et al. (1999). In the M arcus et al.’s experim ent, 7-m onth old infants w ere habituated to sequences of syllables of the form A BA or A BB (e.g., “w e di w e” or “le di di”). M arcus et al. found that infants show ed an attentional preference for novel test sequences of syllables (w hich w e call “sentences”), w hich differed from the habituation stim uli 1 . M arcus et al. argue that the reason for this behavior is the fact that infants extracted “algebra-like rules that represent relationships between placeholders (variables)” (1999). They also concluded that sim ple recurrent networks (and, in general, all netw orks w hose training is based on backpropagation of error) were not able to display this kind of behavior because they could not generalize outside the training space. The issue of generalization outside the training space w as previously addressed in N iklasson and van G elder (1994), and M arcus (1998). In essence, the training space represents the n-dim ensional hyperplane delim ited by the set of training vectors. W e say that a connectionist m odel generalizes to novel stim uli w hen correct output is reliably produced for an input item that For exam ple, after habituated to A BA sequences, the infants spent m ore tim e recognizing novel test sequences of the form A BB than did for A BA sequences, and vice versa. w as not included in the training set (i.e., the netw ork w as never trained on that stim ulus in any position w ithin its input layer). M arcus m aintains that a neural netw ork trained w ith the backpropagation algorithm (or any variant of it) is not able to display such a behavior, because the innate structure of the backpropagation algorithm 2 precludes the netw ork from generalizing to nodes that have not been specifically trained. Elm an agrees that the M arcus experim ent does “indicate that infants discrim inated the difference between the tw o types of sequences” (1999), but he believes that this result m ay be explained by the relationship between the last tw o syllables: infants were able to distinguish that in one case the last two syllables w ere identical (A BB), and in the other case the last tw o syllables w ere different (A BA ). M oreover, Elm an m aintains that it is feasible for a sim ple recurrent netw ork to perform this sam e task, provided the netw ork is presented w ith the sam e background know ledge as infants have (in particular, an exposure to a wide range of syllables that infants have before participating in the experim ent). H aving said that, Elm an perform s an experim ent involving an SRN that aim s to sim ulate the M arcus et al.’s experim ent. There are three phases in Elm an’s sim ulation: 1) the pre-training period, corresponding to the prior experience of the infants in learning to recognize syllables; 2) a second phase corresponding to the habituation task that infants encountered (presenting A BA and A BB sentences); 3) a testing phase involving novel stim uli, as in the infants’ experim ent. A t the end of his sim ulation, Elm an concludes that his results “clearly indicate that the netw ork learned to extend the A BA vs. A BB generalization to novel stim uli” (1999). G ranting Elm an’s basic assum ptions, w e constructed an experim ent that m im ics his sim ulation. W e did not The w eights connecting a given output node are trained independently of the w eights connecting any other output node. Consequently, the set of w eights connecting one output unit to its input units is entirely independent of the set of w eights feeding all other output units. This is called input- output independence, and it is believed to be the m ajor w eak point of backpropagation neural netw orks. It is less clear that the problem arises for com petitive learning netw orks, how ever. See H adley et al (1998) for details.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,002 | 0,010 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,001 | 0,004 |
| Science ouverte | 0,001 | 0,002 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,006 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».