Quantified Grapho-Phonemic Systematicity in Korean Hangeul
Bibliographic record
Abstract
Hangeul, the Korean orthography is well known for its scientific design that emphasizes the link between sounds and letter shapes. However, it hasn’t been asked so far ‘how systematic’ it is. We quantify, for the first time, the grapho-phonemic systematicity of hangeul. We defined Korean phonemes as binary vectors according to articulatory features and then measured the pairwise phonemic distance between phonemes using multiple methods. We measured the pairwise visual distance between letter shapes by (a) stroke share rate, which reflects the original principles of hangeul’s creation, and (b) Hausdorff distance (Huttenlocher et al., 1993), which measures topological difference between images. We then tested the correlation between the phonological distances and the corresponding orthographical distances. Positive correlations clearly indicated that similar letters tend to have similar pronunciations in Korean hangeul. Stroke share rate maximizes hangeul’s grapho-phonemic systematicity. Hausdorff distance, an initial step in the detailed quantifying of visual distance, allows similar calculations to be carried out with any hangeul font and with any other orthography (Jee, Tamariz, & Shillcock, 2021; 2022a; 2022b). Consciously designed to be phonologically transparent, hangeul can be considered as the gold standard of grapho-phonemic systematicity. We discuss the implications of this systematicity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".