MétaCan
Menu
Back to cohort

Does Correct Answer Distribution Influence Student Choices When Writing Multiple Choice Examinations?

2017· article· en· W2596780180 on OpenAlexafffundvenue
Jacqueline Carnegie

Bibliographic record

VenueThe Canadian Journal for the Scholarship of Teaching and Learning · 2017
Typearticle
Languageen
FieldSocial Sciences
TopicEducational Assessment and Pedagogy
Canadian institutionsUniversity of Ottawa
FundersMcMaster UniversityUniversity of Ottawa
KeywordsSummative assessmentMultiple choiceCheatingMathematics educationSelection (genetic algorithm)Computer sciencePsychologyFormative assessmentMathematicsStatisticsSocial psychologyArtificial intelligenceSignificant difference

Abstract

fetched live from OpenAlex

Summative evaluation for large classes of first- and second-year undergraduate courses often involves the use of multiple choice question (MCQ) exams in order to provide timely feedback. Several versions of those exams are often prepared via computer-based question scrambling in an effort to deter cheating. An important parameter to consider when preparing multiple exam versions is that they must be equivalent in their assessment of student knowledge. This project investigated a possible influence of correct answer organization on student answer selection when writing multiple versions of MCQ exams. The specific question asked was whether the existence of a series of four to five consecutive MCQs in which the same letter represented the correct answer had a detrimental influence on a student’s ability to continue to select the correct answer as he/she moved through that series. Student outcomes from such exams were compared with results from exams with identical questions but which did not contain such series. These findings were supplemented by student survey data in which students self-assessed the extent to which they paid attention to the distribution of correct answer choices when writing summative exams, both during their initial answer selection and when transferring their answer letters to the Scantron sheet for correction. Despite the fact that more than half of survey respondents indicated that they do make note of answer patterning during exams and that a series of four to five questions with the same letter for the correct answer would encourage many of them to take a second look at their answer choice, the results pertaining to student outcomes suggest that MCQ randomization, even when it does result in short serial arrays of letter-specific correct answers, does not constitute a distraction capable of adversely influencing student performance. Dans les très grandes classes de cours de première et deuxième années, l’évaluation sommative se déroule souvent par le biais d’examens comportant des questions à choix multiples afin de pouvoir donner rapidement les résultats aux étudiants. Plusieurs versions de ces examens sont souvent préparées et les questions sont brouillées par ordinateur pour dissuader la tricherie. Lors de la préparation de plusieurs versions d’un examen à choix multiples, l’un des paramètres importants à prendre en considération est que chaque version doit être semblable aux autres pour évaluer équitablement les connaissances des étudiants. Ce projet a pour but d’examiner l’influence possible de l’organisation des réponses correctes sur le choix des réponses des étudiants lors de la préparation de plusieurs versions d’un examen à choix multiples. La question spécifique qui a été posée était de savoir si l’existence d’une série de quatre ou cinq questions à choix multiples consécutives pour lesquelles la même lettre représentait la bonne réponse pouvait avoir une influence préjudiciable sur l’aptitude des étudiants à continuer à choisir la bonne réponse alors qu’ils progressent d’une question à l’autre dans la même série. Les résultats des étudiants qui passent de tels examens ont été comparés aux résultats obtenus quand les étudiants passent des examens dont les questions sont les mêmes mais qui ne comportent pas de telles séries. Ces résultats ont été enrichis par les réponses à une enquête auprès des étudiants pour laquelle les étudiants ont été auto-évalués concernant la question de savoir s’ils avaient remarqué la répartition des réponses correctes parmi les choix multiples quand ils passaient des examens sommatifs, à la fois au départ, quand ils choisissaient leurs réponses, et ensuite quand ils transféraient les lettres correspondant à leurs réponses sur la feuille Scanton pour la correction. Malgré le fait que plus de la moitié des répondants aient indiqué qu’ils ne font pas attention à la structuration des réponses pendant l’examen et qu’une série de quatre ou cinq questions ayant la même lettre pour la bonne réponse pourrait encourager beaucoup d’entre eux à regarder de plus près leur choix de réponse, la conclusion concernant les résultats obtenus par les étudiants suggère que la randomisation des questions à choix multiples, même quand elle aboutit à des séries de réponses correctes identifiées par la même lettre, ne constitue pas une distraction capable d’influencer négativement le rendement des étudiants.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.010
metaresearch head score (Gemma)0.014
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Science and technology studies, Scholarly communication
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.230
Threshold uncertainty score0.999

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0100.014
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0280.000
Scholarly communication0.0020.001
Open science0.0010.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.066
GPT teacher head0.412
Teacher spread0.346 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations11
Published2017
Admission routes3
Has abstractyes

Explore more

Same venueThe Canadian Journal for the Scholarship of Teaching and LearningSame topicEducational Assessment and PedagogyFrench-language works237,207