First-year university students’ receptive and productive use of academic vocabulary
Bibliographic record
Abstract
The present study explores academic vocabulary knowledge, operationalised through the Academic Word List, among first-year higher education students. Both receptive and productive knowledge and the proportion between the two are examined. Results show that while receptive knowledge is readily acquired by first-year students, productive knowledge lags behind and remains problematic. This entails that receptive knowledge is much larger than productive knowledge, which confirms earlier indications that receptive vocabulary knowledge is larger than productive knowledge for both academic vocabulary (Zhou 2010) and general vocabulary (cf. Laufer 1998, Webb 2008, among others). Furthermore, results reveal that the ratio between receptive and productive knowledge is slightly above 50%, which lends empirical support to previous findings that the ratio between the two aspects of vocabulary knowledge can be anywhere between 50% and 80% (Milton 2009). This finding is extended here to academic vocabulary; complementing Zhou’s (2010) study that investigated the relationship between the two aspects of vocabulary knowledge without examining the ratio between them. On the basis of these results, approaches that could potentially contribute to fostering productive knowledge growth are discussed. Avenues worth exploring to gain further insight into the relationship between receptive and productive knowledge are also suggested.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.009 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".