Examining Course-Level Conceptual Connections Using a Card Sort Task: A Case Study in a First-Year, Interdisciplinary, Earth Science Laboratory Course
Bibliographic record
Abstract
Universities are recognizing the need to prepare graduates to think conceptually and have the ability to take on complex, real-world problems. Strategies to assess conceptual knowledge are limited and often require more time and effort to complete than is accessible for most undergraduate courses. Card sorting is a very broad technique for understanding how people group concepts, but in higher education has typically been used to show a student’s development towards expert-like thinking in a discipline as a whole. However, it typically does not give much insight into how we should change our teaching. In this paper, using the novel setting of two terms of a first-year, earth and ocean science lab that uses problem-based learning (PBL), we show how one can generate a card sort that is built using course learning goals and then use the analysis to make actionable improvements to course instruction. Using a card sort designed so that the expert sort corresponds to learning goals supported by the lab activities, we found that in both offerings of the course students generally moved towards expert-like sorting with a reduction in novice-like sorting. A striking feature stood out in both terms of the course, with one question scoring significantly lower than any other expert pairings, despite a change in the wording of that question between terms. This suggests that our course materials do not promote this specific conceptual connection that we had expected and gives us a clear place to look for issues in our course material. In a broader context, our results suggest that tailoring card sort questions to material at a course level, rather than at the discipline level, can provide a manageable, routine assessment of conceptual knowledge in students, while also providing feedback on the quality of course materials.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.042 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.014 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.013 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".