Looking for Categorical Perception in a Dot-Pattern Classification Task
Bibliographic record
Abstract
Looking for Categorical Perception in a Dot-Pattern Classification Task Gyslain Giguere (giguere.gyslain@courrier.uqam.ca) Universite du Quebec a Montreal, Departement de psychologie, c/o LEINA, C.P. 8888, Succ. CV, Montreal, QC, H3C 3P8, Canada Guy L. Lacroix (guy_lacroix@carleton.ca) Carleton University, Department of Psychology, B550 Loeb Building, 1125 Colonel By Drive, Ottawa, ON, K1S 5B6, Canada Serge Larochelle (serge.larochelle@umontreal.ca) Denis Cousineau (denis.cousineau@umontreal.ca) Universite de Montreal, Departement de psychologie, C.P. 6128, Succ. CV, Montreal, QC, H3C 3J7, Canada category learning. That is, they naturally emerged from the probabilistic distortion creation technique. Defining categorization Category learning entails two processes: similarity-based clustering, which involves positioning objects in a multi- dimensional psychological space, and labeling, which involves associating arbitrary linguistic labels with each acquired cluster. In human cognition, categorization serves an optimization purpose; a way of overcoming limited processing resources via the reduction of information. One such optimization procedure is learned categorical perception (LCP). It maximizes categorical knowledge by enhancing within- category similarity and/or reducing between-category similarity. While often taken for granted, LCP has yet to be shown convincingly in empirical work. A classic categorization task is the dot-pattern classification paradigm. Participants must learn to categorize exemplars created by probabilistically distorting prototypical patterns. This technique is widely used, because the properties of these artificial categories are thought to resemble those of real-world, natural ones. (Homa, 1984). We hypothesized that if dot-patterns are representative of real-life categories, and categorical perception is an optimal way of enhancing information use, then LCP should be found in a dot-pattern task. Figure 1: Mean inter-stimulus similarity scores for participants who did or did not categorize before judging. Discussion LCP was not found in this experiment. Rather, the results suggest that the dot-pattern classification paradigm entails the labeling process only. Hence, it may be argued that the dot-pattern classification task is not useful to understand similarity-based clustering. Acknowledgments Our experiment The methodology was based on Shin and Nosofsky’s (1992) Experiment 1. Half of our participants were asked to make similarity judgments about pairs of never before seen dot- patterns, while the other half was asked to categorize these exemplars for 15 blocks before making similarity judgments. This research was made possible by grants awarded by the Universite de Montreal, the Natural Sciences and Engineering Council of Canada, the Fonds Quebecois de Recherche sur la Nature et les Technologies and the Fonds de Recherche en Sante du Quebec. References Homa, D. (1984). On the nature of categories. Psychology of Learning and Motivation, 18, 49-94. Shin, H.J., & Nosofsky, R.M. (1992). Similarity-scaling studies of dot-pattern classification and recognition. Journal of Experimental Psychology: General, 121, 278-304. Results As seen in Figure 1, training with dot-pattern categories did not modify inter-stimulus similarities. When exploring the similarity data using MDS, we discovered that the expected result was not found because the clusters existed before
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.022 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.006 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.019 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".