The very first sentence in research article introductions: A rhetoric comparative approach
Bibliographic record
Abstract
The current study explores the rhetoric and stylistic properties of the very first sentence that scholars generate in their research article introductions. The study draws upon a corpus of 502 sentences written in the fields of linguistics and translation, half of which are collected from national low-impact journals affiliated with Gulf universities in the Middle East while the other half are elicited from international high-impact journals. The study shows that half of the authors in high-impact journals as opposed to a quarter of the authors in low-impact journals provide citations to their very first sentence. These preferences are accounted for by the distinction drawn by Swales (1990) between centrality claims and topic generalizations under Move 1. Contra the predictions made by Create A Research Space Model proposed by Swales (1990, 2004), the results show that the authors of high-impact journals are more liberal in starting their introduction with a sentence of Move 2 or 3 type. In contrast, the authors of low-impact journals prefer to begin with a sentence of Move 1 type that is shorter in word count, more metaphorical, less academic as well as full of typos and grammatical errors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.035 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.007 | 0.004 |
| Science and technology studies | 0.003 | 0.005 |
| Scholarly communication | 0.005 | 0.006 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".