Academic Vocabulary Use in Doctoral Theses: A Corpus-Based Lexical Analysis of Academic Word List (AWL) in Major Scientific Disciplinary Groups
Bibliographic record
Abstract
Since the development of academic word list (AWL) by Coxhead (2000), multiple studies have attempted to investigate its effectiveness and relevance of the included academic vocabulary in the texts or corpora of various academic fields, disciplines, subjects and also in multiple academic genres and registers. Similarly, this study also aims at investigating the text coverage of Coxhead’s (2000) AWL in Pakistani doctoral theses of two major scientific disciplinary groups (Biological & health sciences as well as Physical sciences); furthermore the study also analyses the frequency of the AWL word families to extract the most frequent word families in the theses texts. In order to achieve this goal, a pre-built corpus of Pakistani doctoral theses (PAKDTh) (Aziz, 2016) comprises of 200 doctoral theses from two major scientific disciplinary groups was used as textual data. Using concordance software AntConc version 3.4.4 (Anthony, 2016), computer-driven data analysis revealed that in total 8.76% (496839 words) of the text in Pakistani doctoral thesis corpus is covered by the AWL words. Further distributing the analysis per sub-lists, shows that the first three sub-lists of AWL accounted for almost 57% of the whole text coverage. An attempt was made to further analyze the AWL text coverage by considering the frequency of occurrences in terms of word families. The findings showed that among 570- word families of Coxhead’s (2000) AWL, 550-word families with the sum of 96.49% are found to occur more than 10 times in PAKDTh corpus, which are taken as word families used in the corpus. This study concludes that Coxhead’s (2000) AWL is proved effective for the writing of theses. On the basis of the findings, further possible academic implications are discussed in detail.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.014 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.011 | 0.011 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.000 | 0.003 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".