Old ‘truths’, new corpora: revisiting the word order of conjunct clauses in Old English
Bibliographic record
Abstract
In Bech (2001a, 2001b), I took issue with the oft-repeated claim that Old English conjunct main clauses are commonly verb-final, and disproved it. However, the myth persists. In the meantime, theYork–Toronto–Helsinki Parsed Corpus of Old English Prose(YCOE, Tayloret al.2003) has been created, so the time has come to revisit this topic and consider it in light of new, extensive and generally accessible data. Using the YCOE corpus, I confirm and expand on Bech's (2001a, 2001b) empirical findings, showing that (i) OE conjunct clauses are neither typically verb-final nor verb-late, but they are more frequently verb-final and verb-late than non-conjunct clauses are; and (ii) verb-final and verb-late clauses are typically conjunct clauses. These two perspectives must be kept apart: in the first, the starting point is the entire body of conjunct clauses, and in the second it is the entire body of verb-final/verb-late clauses. I propose that the failure to distinguish between the two perspectives, i.e. whether it is conjunct clauses or word order that constitutes the point of departure, is the origin of the misconception concerning conjunct clauses and word order. In order to establish whether this distinction has been fuzzy all along, or whether it must be ascribed to distorted referencing in the course of a century of research, I trace the research on this topic back to the end of the nineteenth century. I show that the alleged verb-finality of conjunct clauses may be ascribed to awhisper-down-the-laneeffect – the retelling of the story has changed the story.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.044 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.006 | 0.007 |
| Science and technology studies | 0.003 | 0.007 |
| Scholarly communication | 0.005 | 0.014 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.007 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".