English professors, nursing students, and the HESI A2 Exam
Bibliographic record
Abstract
As English professors, we never thought we would become immersed in the world of nursing and its associated Health Education Systems Inc.Admission Assessment (commonly known as HESI or, more formally, HESI A2) standardized test.That changed one day in 2017 when a former 1styear composition student stopped by my (Jennifer's) office, clutching a book, and asked, "Can you help me prepare for my nursing entrance exam?"When I demurred, explaining my expertise was in English, the student replied, "but it's all English."And so it began.That day, I learned that this exam has three English sections, in addition to science and math.My student wasn't concerned about the latter subjects.But Englishgrammar, in particular-was rapidly looking like a barrier to her goals.With two failed attempts under her belt and two attempts remaining, this hard-working young person was not prepared to give up.As I perused her prep book-direct from the test makers-my initial bemusement turned to, if I'm being honest, a bit of anger.Of the 10 practice questions, one had no correct answer choice (a deep dive on the associated website showed a correction, but it took a fathoms-deep dive to find it).Some of the questions seemed reasonable, but others seemed advanced for native speakers, let alone English-language learners (ELLs).At one point, I stepped out of the meeting with the student to show Heather, a colleague, for a sanity check.I wasn't crazy.The test was crazy hard for ELLs.That day, unbeknownst to me at the time, an unlikely partnership formed in the realm of test-prep. PARTNERSHIP IN ITS INFANCY Jennifer's perspectiveThat first year, the student and I stumbled our way through (or at least I stumbled), with me making practice questions and explaining concepts.Her third attempt showed improvement, but-even better-she now had the language (and comfort level) to tell me "the questions aren't like this one.They're more like this." Thank goodness for that!Each redirection helped me feel like I was stumbling a little less and that I could find a way to help her.With her experience taking the test and more practice test books, we worked our way through everything from homophones to predicate nominatives.I became more confident, and so did the student.Our first rule became "rule out two answers and have reasons."We met for
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gpt | no category Domain: not available · Genre: Other About the Canadian research system: no · About a Canadian topic: no | Other design | low |
| grok | no category Domain: not available · Genre: Other About the Canadian research system: no · About a Canadian topic: no | Other design | low |
| opus | no category Domain: not available · Genre: Other About the Canadian research system: no · About a Canadian topic: no | Other design | low |
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 3 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".