Understanding the interconnection between assessment of student learning and generative artificial intelligence
Bibliographic record
Abstract
This special issue provides a glimpse into the questions and activities that higher education institutions around the globe are actively pursuing to address the interconnectedness between generative artificial intelligence (GenAI) and the assessment of student learning. The eight papers in this special issue come from six countries, illustrating that in every corner of the world, the interconnection between student assessment and GenAI is an urgent question. How is GenAI impacting the way educators do assessment? What can educators change in the process of assessment? How can educators benefit from GenAI to prepare the next generation? We are publishing this special issue to engage in the conversation about who is doing what and what is being considered, theorized, examined, and researched. The papers have been categorized into three groupings: conceptual ideas, lesson applications, and research studies. There is a constant theme of practicality and usefulness in all eight papers. In short, every paper has as its focus the provision of thoughtful, ethically focused help in the questions that are asked, the advice given, the applications provided, and the cautions and suggestions noted. We suspect that readers will be inspired to roll up their sleeves too, as higher education collectively considers how to reimagine assessment of learning in an AI-driven world. The issue begins with Cesare Giulo Ardito providing an inspiring conceptual examination of the effectiveness, vulnerabilities, and ethical implications of AI detection tools in the context of preserving academic integrity. In the next paper, Shamola Pramjeeth and Priya Ramgovind move from conceptual questions to a study revealing the current perceptions of academics and teaching and learning specialists on using AI to reconceptualize assessment and teaching and learning structures. After discussing concerns and identifying common perceptions, the issue moves on to explore the ethical applications of GenAI tools. Christopher Hill and Jace Hargis present a framework for a versatile 3-h module applicable to various disciplines. The module provides a concrete example of how to engage students in a discussion about the responsible use of GenAI in assignments and how to foster a dialogue between students and faculty on crafting effective policies for GenAI utilization. Daniel Dale provides an example of an assessment that was used in a composition class. Daniel hopes to “show how composition studies can provide a useful framework for thinking about integrating generative AI assignments into all courses.” Next, Christian Coenen explores how AI was trained to provide personalized comments to students in the instructor's style in high enrollment courses. The article details the innovation, implementation, and context of this approach, reflecting on the instructor's experience and analyzing its effectiveness and student learning impacts. Using a disability justice approach, Azeezah Jafry and Jessica Vorstermans guide readers through a reflection of how GenAI might be a generative invitation to engage more fully in differentiated assessments and a renewed commitment to access for all students. Azeezah Jafry and Jessica Vorstermans highlight the need for institutions “to meet the new shift to GenAI in ways that are non-punitive, rooted in access, and prepare students to enter the world we have today.” Institutional use of AI is the focus of the last two papers. David DiSabito, Lisa Hansen, Thomas Mennella, and Josephine Rodriguez explore the potential of GenAI to streamline the assessment process, making it more efficient, equitable, and objective through the development of a proprietary GenAI tool called Walter. The study considers GenAI's role in assessing student evidence against human assessment while addressing challenges such as data collection, coordination, and the need for well-defined and precise rubrics. Ruth Slotnick and Joanna Boeing, in the final paper in the issue, study the potential of GenAI to enhance qualitative research in higher education assessment. They examined the use of two large language models, Google's Bard and OpenAI's ChatGPT, to analyze qualitative data focusing on diversity, equity, and inclusion in program assessment, assessing the strengths and limitations of each large language model. Given the massive financial and human investment in AI, we expect to see ever more interest, reflection, and research about the relationship between the assessment of student learning and artificial intelligence. This issue provides a window into what higher education colleagues are asking, studying, and developing. We hope it will inspire you with innovative ideas. The authors in this special issue all identify challenges to the use of AI, and in so doing provide several new and exciting opportunities for future theoretical and experiential development. We hope you will find interesting avenues for future research and application in these papers. Dr. Eliana El Khoury is a recognized leader in the field of alternative assessment, holding a PhD from the University of Calgary. She serves as an assistant professor at Athabasca University, where she works closely with experts across different disciplines to improve how students are evaluated. Her primary focus is on ensuring quality and fairness in educational assessments, and she actively incorporates new ideas and technologies to achieve this goal. Dr. El Khoury is the founder and chair of the Symposium on Alternative Assessment, a key event that brings together specialists to discuss and develop more effective assessment methods. Her efforts in the symposium and beyond aim to make assessment tools that better reflect student learning and support educational success. Dr. El Khoury's work is dedicated to enhancing educational practices through more thoughtful and supportive assessment strategies. Deborah Homuth has been an educator for many years in avariety of roles including as an elementary and secondary teacher, curriculumconsultant, and secondary school principal. She built Desire2Learn's VirtualHigh School (Ontario) a fully online Ontario secondary school. She was recruited by Pearson as the National LearningTechnology Consultant for Canada assisting provincial and territorialMinistries and Departments of Education across the country and seniormanagement at the school district level to envision and then implement 21stcentury learning enabled by technology. She was responsible for reviewing thenewest educational technology being built or acquired by Pearson worldwide todetermine its applicability to education in Canada. She led RFP development, wrote several whitepapers about educational technology, co-chaired a nationaleducation conference of educational leaders focused on making the shift tofully online and blended learning JK to post-secondary, and providedprofessional learning on reimagined assessment in Canada. As Senior Manager andDirector of Curriculum for TVO, Ontario's publicly funded broadcaster andmanager of the Independent Learning Centre, Deborah was responsible forshifting the ILC from a correspondence model to TVO's fully web-based school fromvision to launch serving 25,000 + students annually in English and French. Shewas Senior Director of H2Learning Consultants responsible for conductingnational needs assessments for educational providers and for designing onlinelearning for post-secondary institutions. As Senior Manager at ConestogaCollege in Ontario she led the development of new online programs leading todiplomas and degrees. Deb is currently a doctoral student in distance education.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".