Bibliographic record
Abstract
Discourse connectives (DCs) are multi-functional devices used to connect discourse segments and fulfill interpersonal levels of discourse. This study investigates the use of selected 80 DCs within 11 categories in the argumentative essays produced by L1 and L2 university students. The analysis is based on the International Corpus Network of Asian Learners of English (ICNALE) which consists of essays written by native speakers (NS) and non-native speakers (NNS) from 10 countries and regions in Asia. WordSmith Tools were used to generate the quantitative profile of the DCs, while follow-up qualitative analysis in the context of usage provided additional interpretive insights. The total frequency of DCs used by Hong Kong and Singaporean students is significantly less than do L1 writers, mainly because the addictive and is by far less frequent in the essays produced by L2 writers. Hong Kong students use much more enumerating, resultive and summative DCs than both L1 writers and L2 writers from Thailand and Singapore. Thai students, on the other hand, employ the causal device because much more than both L1 and other L2 writers. Hong Kong and Singaporean students are more formal in tone than L1 and Thai students when using the adversative and resultive DCs. Despite the apparent differences, there are considerable similarities of usage, with and, but, because, so, however and therefore occurring among the top 10 most frequently used devices of both L1 and L2 writers, although with strikingly different frequencies. These findings shed light on the pragmatic uses of DCs by L1 and L2 writers as a way to influence the interpretation of the message, and thus succeed in achieving their communicative intentions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.027 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.005 | 0.003 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.006 | 0.003 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".