MétaCan
Menu
Back to cohort
Record W4414203932 · doi:10.1145/3767736

Can Alternative Grading Improve Student Interactions in Automatically Graded Programming Assignments?

2025· article· en· W4414203932 on OpenAlexaff
Katharine Kerr, N Bradley, Reid Holmes

Bibliographic record

VenueACM Transactions on Computing Education · 2025
Typearticle
Languageen
FieldComputer Science
TopicTeaching and Learning Programming
Canadian institutionsUniversity of British Columbia
Fundersnot available
KeywordsGrading (engineering)DebuggingOracleFormative assessmentSuiteTest suiteUnit testingSoftwareSummative assessmentSoftware development

Abstract

fetched live from OpenAlex

Background. Automated assessments, often implemented as test-based autograders, are widely used in the form of hidden oracles, providing automated formative feedback to students on their code solutions. While autograders provide some educational benefits—particularly around efficient scaling of grading—they have been shown to encourage adverse student behaviors that hinder learning. For example, students try to debug their solutions into existence through trial-and-error debugging in pursuit of maximum points, without the valuable, careful introspection on their work. Objectives. This study investigates how a grade scale affects students’ software development behaviors in response to automated feedback. In contrast to an autograder with an oracle suite of test cases, industrial developers do not have access to an oracle test suite that can tell them what cases their code handles or mishandles. Developers must instead rely on careful reasoning to co-evolve their test code with their product code to produce and maintain high-quality software. We hypothesize that alternative grading can be an effective tool to influence students to practice more reflection during development while maintaining the infrastructural and pedagogical benefits of automated assessments. Methods. We deployed a coarse-grained, alternative grading approach—bucket grading—which assessed solutions to be in one of only four bucket grades, representing a high-level assessment of the student’s project quality, to a third-year post-secondary, project-based software engineering class with 300+ students. This study uses a mixed qualitative and quantitative methodology to compare a previous offering of the course that used a traditional, points-based grading scheme against the bucket grading offering. Findings. We find that coarse-grained formative feedback via bucket grading improves student-autograder interaction: students wrote stronger, more focused test suites because they reflected more deeply before making changes. Students were appreciative of bucket grading since it provided them leniency in their grade and greater clarity on the overall quality of their solutions. Ultimately, we will continue to use this approach going forward since incorporating bucket grading meaningfully improved both the staff and student autograder experiences.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.022
metaresearch head score (Gemma)0.144
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.022
Threshold uncertainty score0.115

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0220.144
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.001
Science and technology studies0.0010.001
Scholarly communication0.0050.004
Open science0.0020.003
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0080.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.019
GPT teacher head0.347
Teacher spread0.328 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueACM Transactions on Computing EducationSame topicTeaching and Learning ProgrammingFrench-language works237,207