Vigilance Procedure Generalization for Recurrent Associative Memories
Bibliographic record
Abstract
Vigilance Procedure Generalization for Recurrent Associative Memories Sylvain Chartier (sylvain.chartier@courrier.uqam.ca) Centre de recherche de l’Institut Philippe Pinel de Montreal 10,905 Henri-Bourassa Est, Montreal, QC, H1C 1H1, Canada Sebastien Helie (helie.sebastien@courrier.uqam.ca) 1 , Robert Proulx (proulx.robert@uqam.ca) 2 Mounir Boukadoum (boukadoum.mounir@uqam.ca) 1 Departement d’informatique, Universite du Quebec a Montreal, Departement de psychologie, Universite du Quebec a Montreal POB 8888, station Downtown, Montreal, QC, H3C 3P8, Canada Introduction In our ever changing world, each experienced stimulus differs from the previous. This variation can be explained using two sources: signal noise and exemplars. To overcome the possibly infinite number of stimuli, humans are able to group these unique stimuli into a finite number of categories. In particular, human cognition enables adaptation in many environments, which necessitate a broad range of behaviors which is a function of context. Most unsupervised neural networks cannot deal with such variability. One exception is the family of ART networks, which were proposed to solve the stability - plasticity dilemma (e.g. Carpenter & Grossberg, 1987). These models are able to achieve the desired behavior by using a vigilance procedure. However, this procedure has never been generalized to other classes of unsupervised neural networks, in particular recurrent associative memories. This study proposes a generalization of the vigilance procedure that can be implemented from one-shot binary input learning models (e.g. Hopfield, 1982) to iterative learning real-value patterns models (e.g. Chartier & Proulx, 2005). Vigilance Procedure The role of vigilance is to specify whether a novel stimulus belongs to a previously learned category or a new one. To accomplish this, a new stimulus is shown to the network, and it iterates until convergence. The resulting stable state is compared with the initial stimulus using standard correlation: if the correlation between an initial input (x(0)) and its corresponding attractor x(c) is lower than the vigilance parameter’s value ( ρ ), the new stimulus forms a new category. On the other hand, if the correlation between the stimulus and its corresponding attractor is higher than the vigilance parameter’s value, the new stimulus is integrated into this existing category. In this case, the new stimulus modifies the position of the attractor by using the following average between the initial input and the attractor. x = z ( α x (0) + x ( c ) ) x (0)(1 − z ) 1 + α z where, x is the network’s state used by the given model’s learning rule, α (0 < α << 1) is a parameter which quantifies the effect of the initial input in x and z return 1 if the correlation is greater that ρ and 0 otherwise. Thus, if z = 0, then x = x (0) (initial stimulus); if z = 1, x = ( α x (0) + x (c)) /(1 + α ) (weighted average of the initial and stable states). This procedure is illustrated in Figure 1. Figure 1: Vigilance procedure Conclusion This study shows how to implement a vigilance procedure into RAMs. Consequently, the vigilance procedure is no longer exclusive to competitive networks, which broadens the application domain of RAMs. References Carpenter, G. A. & Grossberg, S. (1987). A massively parallel architecture for a self-organizing neural pattern recognition machine. Computer Vision, Graphics, and Image Processing, 37, 54-115. Hopfield, J. J.(1982). Neural networks and physical systems with emergent collective computational abilities, Proceedings of the Natural Academy of Sciences (U.S.A.), Chartier, S. & Proulx, R., (2005). NDRAM: Nonlinear Dynamic Recurrent Associative Memory for bipolar and non bipolar learning. IEEE Transactions on Neural Networks, 16, 1393-1400.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.027 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.001 | 0.004 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.008 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".