Seeing stems everywhere: Position-independent identification of stem morphemes.
Bibliographic record
Abstract
There is broad consensus that printed complex words are identified on the basis of their constituent morphemes. This fact raises the issue of how the word identification system codes for morpheme position, hence allowing it to distinguish between words like overhang and hangover, and to recognize that preheat is a word, whereas heatpre is not. Recent data have shown that suffixes are identified as morphemes only when they occur at the end of letter strings (Crepaldi, Rastle, & Davis, 2010, "Morphemes in Their Place: Evidence for Position-Specific Identification of Suffixes," Memory & Cognition, 38, 312-321), which supports the general proposal that the word identification system is sensitive to morpheme positional constraints. This proposal leads to the prediction that the identification of free stems should occur in a position-independent fashion, given that free stems can occur anywhere within complex words (e.g., overdress and dresser). In Experiment 1, we show that the rejection time of transposed-constituent pseudocompounds (e.g., moonhoney) is longer than that of matched control nonwords (e.g., moonbasin), suggesting that honey and moon are identified within moonhoney, and that these morpheme representations activate the representation for the word honeymoon. In Experiments 2 and 3, we demonstrate that the masked presentation of transposed-constituent pseudocompounds (e.g., moonhoney) facilitates the identification of compound words (honeymoon). In contrast, monomorphemic control pairs do not produce a similar pattern (i.e., rickmave did not prime maverick), indicating that the effect for moonhoney pairs is genuinely morphological in nature. These results demonstrate that stem representations differ from affix representations in terms of their positional constraints, providing a challenge to all existing theories of morphological processing.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".