Processing Load Imposed by Line Breaks in English Temporal Wh-Questions
Bibliographic record
Abstract
Prosody plays an important role in online sentence processing both explicitly and implicitly. It has been shown that prosodically packaging together parts of a sentence that are interpreted together facilitates processing of the sentence. This applies not only to explicit prosody but also implicit prosody. The present work hypothesizes that a line break in a written text induces an implicit prosodic break, which, in turn, should result in a processing bias for interpreting English wh-questions. Two experiments-one self-paced reading study and one questionnaire study-are reported. Both supported the "line break" hypothesis mentioned above. The results of the self-paced reading experiment showed that unambiguous wh-questions were read faster when the location of line breaks (or frame breaks) matched the scope of a wh-phrase (main or embedded clause) than when they did not. The questionnaire tested sentences with an ambiguous wh-phrase, one that could attach either to the main or the embedded clause. These sentences were interpreted as attaching to the main clause more often than to the embedded clause when a line break appeared after the main verb, but not when it appeared after the embedded verb.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.017 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".