MétaCan
Menu
← Back to cohort

The Brave New World of Anatomy: Using AI to Grade Practical Examinations

2021· article· en· W3172612656 on OpenAlexaff
Alex B. Bak, J Bernard, Anthony N. Saraco, Josh Mitchell, Ilana Bayer, Ranil Sonnadara, Bruce Wainman

Bibliographic record

VenueThe FASEB Journal · 2021
Typearticle
Languageen
FieldMedicine
TopicInnovations in Medical Education
Canadian institutionsHamilton Health SciencesMcMaster UniversityUniversity of Toronto
Fundersnot available
KeywordsGrading (engineering)Computer scienceCorrectnessArtificial intelligenceSet (abstract data type)SpellSyllabusCurriculumNatural language processingMathematics educationPsychology

Abstract

fetched live from OpenAlex

In a shift of medical education to a competency‐based curriculum, practical examinations (PEs) are an effective but resource‐intensive method of evaluating anatomy students. The short answer format of PEs requires evaluators familiar with the content to mark the exams. Moreover, the increasing transition to online anatomy courses could result in students losing the PE practice they would receive during in‐person sessions. By virtue of the technical and close‐ended nature of typical PE answers as well as grading usually being a binary ‘correct’ or ‘incorrect’ classification with no partial credit, it was hypothesized that the limited lexicon would allow for accurate grading using artificial intelligence disciplinessuch as natural language processing and decision trees (DTs). This research was done as a first step towards making an intelligent tutoring system for anatomy students. The study used the winter semester online PE results (n = 371) from McMaster University Faculty of Health Sciences’ anatomy and physiology course as the data set. For each of the 54 questions, a 10‐fold cross‐validation process was used where 90% of the answers (training set) trained the DT. After removing common words unrelated to correctness (“the”, “a”, “an”, etc.), each DT was comprised of unique words that appeared in student answers in a tree‐like structure of nodes. Each node has an associated word as well as a correct/incorrect classification label and splits into sub‐nodes (creating the tree‐like structure). The remaining 10% of the answers (testing set), was marked by the generated DTs by traversing the tree starting from the top‐most node. After traversing the tree, the classification label of the final node became the grade for the student's answer. Accuracy for each question was calculated as the number of proper classifications by the algorithm over the total number of answers. When the answer marked by the DT were compared to the answers marked by staff and faculty, the DT achieved an average of 94.49% accuracy in grading every non‐blank student answer across all 54 questions. It was found that accuracy was negatively correlated to the number of unique words in the set of answers (‐0.71, p<0.07), which was consistent with the initial hypothesis. As features such as spellchecking were not included in the algorithm to reduce the number of variables, the current results may underestimate the effectiveness of automated PE grading by DTs. The accuracy attained by the algorithms suggests that machine learning algorithms such as NLP and DTs may be used to reduce the workload of manual PE grading by instructional staff and mark a step towards developing an intelligent online PE tutoring system for anatomy.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.009
metaresearch head score (Gemma)0.052
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.009
Threshold uncertainty score0.047

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0090.052
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0030.002
Science and technology studies0.0010.001
Scholarly communication0.0040.004
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.062
GPT teacher head0.420
Teacher spread0.359 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2021
Admission routes1
Has abstractyes

Explore more

Same venueThe FASEB Journal→Same topicInnovations in Medical Education→French-language works237,207→