What colours a letter? The deep learned structure of synaesthesia in two linguistic groups
Bibliographic record
Abstract
The development of grapheme-colour synaesthesia is a slow process, in which children gradually develop a consistent set of colours associated with each letter of the alphabet, beginning before age 6 and continuing until at least age 12 (Simner & Bain, 2013). During this time, a wide range of factors influence the specific colour associated with a given letter, factors that recent research is beginning to tease apart. These include semantic associations with letters, their order in the alphabet, phonological characteristics, shape, and frequency of use (e.g. Simner et al., 2005; Asano & Yokosawa, 2013; Watson, Akins & Enns, 2012). Here we directly compare these influences and several never previously-described in two distinct populations of synaesthetes (native Czech and English speakers). We find that while some factors (e.g. letter shape and alphabetical order) have similar effects in the two groups, other factors are particular to each language. For instance, the regular relationship between phonemes and graphemes in Czech enables the phonological similarity of letters to influence their colours, unlike in English, and Czech has several diacritical marks that influence the colour of their letters in different ways. The overall picture that emerges is one of complex interacting influences competing with each other over the course of synaesthetic development, leading to the apparently idiosyncratic sets of colours of adult synaesthetes. Meeting abstract presented at VSS 2014
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.000 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".