The rise of<i>it</i>-clefting in English: areal-typological and contact-linguistic considerations
Bibliographic record
Abstract
Recent areal and typological research has brought to light several syntactic features which English shares with the Celtic languages as well as some of its neighbouring western European languages, but not with (all of) its Germanic sister languages, especially German. This study focuses on one of them, viz. the so-calledit-cleft construction. What makes theit-cleft construction particularly interesting from an areal and typological point of view is the fact that, although it does not belong to the defining features of so-called Standard Average European (SAE), it has a strong presence in French, which is in the ‘nucleus’ of languages forming SAE alongside Dutch, German, and (northern dialects of) Italian. In German, however, clefting has remained a marginal option, not to mention most of the eastern European languages which hardly make use of clefting at all. This division in itself prompts the question of some kind of a historical-linguistic connection between the Celtic languages (both Insular and Continental), English, and French (or, more widely, Romance languages). Before tackling that question, one has to establish whetherit-clefting is part of Old (and Middle) English grammar, and if so, to what extent it is used in these periods. In the first part of this article (sections 2 and 3), I trace the emergence ofit-clefts on the basis of data fromThe York–Toronto–Helsinki Corpus of Old English ProseandThe Penn–Helsinki Parsed Corpus of Middle English, second edition. Having established the gradually increasing use ofit-clefts from OE to ME, I move on to discuss the areal distribution of clefting among European languages and its typological implications (section 4). This paves the way for a discussion of the possible role played by language contacts, and especially those with the Celtic languages, in the emergence ofit-clefting in English (section 5). It is argued that contacts with the Celtic languages provide the most plausible explanation for the development of this feature of English. This conclusion is supported by the chronological precedence of the cleft construction in the Celtic languages, its prominence in modern-period ‘Celtic Englishes’, and close parallels between English and the Celtic languages with respect to several other syntactic features.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.002 | 0.005 |
| Scholarly communication | 0.003 | 0.005 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".