Investigating the Construct of a Listening Test to Assess Pilots’ Comprehension: a Step-by-Step Project for Test Developers
Bibliographic record
Abstract
Pilots and air traffic controllers must demonstrate their ability to listen and speak the language used in radiotelephony communications demonstrated by completing a language test. In this context, it is crucial to assess both interactive listening, when listening occurs together with speaking, and listening in isolation, when there is no speaking or interaction. The purpose of assessing listening in isolation is to reduce the influence of skills that are not relevant to the construct, that is, to minimize construct irrelevant variance (S. Messick 1994). This article describes a project that can be followed by test developers to address the initial step in the development of a test to assess pilots’ listening in isolation: the construct definition. The project is framed within an interactionalist perspective wherein a test construct is defined based on a combination of the abilities that those taking the test should have and the tasks that they should be able to perform (L. Bachman 2007). It is also informed by the work of L. Bachman/ A. Palmer (2010) and the framework proposed by U. Knock/ S. Macqueen (2020) for the development of language assessments for professional purposes. The project outlined in this article may also be of interest to test developers who wish to investigate different constructs of aeronautical English tests, as well as those involved in the development of other types of language assessments for professional purposes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.102 | 0.103 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.005 | 0.005 |
| Open science | 0.004 | 0.008 |
| Research integrity | 0.002 | 0.006 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".