Transparent Assessments: The Potential Benefits of Training Students in the How and Why of Evaluation
Bibliographic record
Abstract
The ideal assessment accurately measures student knowledge and skills without causing debilitating anxiety. This study investigates the impact of two interventions on student anxiety, perceptions, and performance: increased transparency in evaluative techniques prior to the assessment, and working with peers during the assessment. Evaluation anxiety can reflect discrepancies between self‐perceived assignment quality and the received mark, which may reflect instructors' use of tacit (i.e. knowing good work when I see it) as well as explicit knowledge (e.g. marking schemes). The impact of increased knowledge transfer from instructor to students was examined in an upper year Physiology seminar class using concept maps as summative evaluations. Students engaged in a training procedure involving rubric presentation (explicit knowledge transfer), small group evaluations of exemplar concept maps, and an instructor‐led discussion of the exemplars (tacit knowledge transfer). Students found the instructor‐led discussion the most helpful of the three steps. Students consistently ranked concept maps as better performance indicators than multiple choice exams, but results varied as to their anxiety‐provoking potential. Working in groups was only deemed useful if students already had a preliminary concept map to discuss. In comparison to students in previous classes who were not provided explicit training, students reported more positive perceptions towards concept mapping and generated higher quality maps. The second study involved a large, inquiry‐based, upper year Physiology course using homework, group quizzes, and individual multiple choice and short answer exams. Since these evaluative techniques are relatively straightforward, we focused on differences between individual and group assessments. Group quizzes invoked less anxiety than individual exams in all students, but the difference was greater in self‐reported B students than in self‐reported A students. Yet, B students considered these assessments poorer indicators of their knowledge than A students. Comments indicated that A students worry about “free‐loaders”, and B students often feel that quiz results reflect the smartest student's knowledge, not their own. Explicit instructor explanations of the formative advantages of group quizzes, such as self‐elucidation for stronger students and peer teaching for weaker students, might increase student buy‐in. These results highlight the importance of explicit instructor explanations of how and why instructors evaluate student work. Support or Funding Information This work was supported in part by the Senate Research Committee of Bishop's University. This abstract is from the Experimental Biology 2018 Meeting. There is no full text article associated with this abstract published in The FASEB Journal .
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.023 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".