Bibliographic record
Abstract
I have been a professional historian for more than four decades, and I never once envied physicists — at least not until I encountered the “Force Concept Inventory” (Hestenes, Wells, & Swackhamer, 1992). This set of questions allows teachers of physics to determine whether students have internalized the basic Newtonian model taught in physics courses or whether they automatically fall back on “Aristotelian” notions of objects moving in space. It gives instructors in the field a simple instrument for determining the sophistication of students entering their classes and for evaluating the success of their own teaching strategies over a semester. This kind of disciplinary consensus about learning goals allows for the kind of far reaching and impressive assessment of learning that has allowed scholars of teaching and learning like Richard Hake to make convincing claims about the relative value of different teaching strategies (Hake, 1998; Bain, 2004). A historian reading such work is apt to be immediately struck by the absence in his or her own field of this kind of agreement about what should be taught and what constitutes reasonable evidence that it has been learned. While physicists certainly argue about theoretical issues at the forefront of knowledge, they do not need to spend a great deal of effort justifying either the truth value of Newtonian mechanics or its relevance to the curriculum. In a discipline such as history, by contrast, undergraduates can enter contested spaces from their first day in the college classroom, and the subject matter that they are studying is as varied as the cultures that have left a trace on this planet. Both the ambiguity of sources and the co-existence of mutually contradictory interpretations would seem to dictate that history is and is apt to remain a “fuzzy” discipline. This relative dearth of consensus is a result of the nature of the phenomena historians study, rather than any great deficiency in their profession, but it does lead to difficulties in creating a credible scholarship of teaching and learning. In a field in which reasoning is, of necessity, somewhat nebulous, it can be daunting to develop a clear consensus on what constitutes evidence that learning has occurred. Yet, assessment is at the core of the entire SoTL enterprise. It is difficult to imagine a robust scholarship of teaching and learning unless our work is cumulative and built on previous research and unless there is a means to systematically evaluate the validity of claims being made about student learning. In the now canonical formulation of Lee Shulman, the scholarship of teaching and learning must be “public, susceptible to critical review and evaluation, and accessible for exchange and use by other members of one’s scholarly community” (Shulman, 1998). Historians have made considerable progress in making SoTL public and accessible. An international society of historians working in the field has been created with its own website and newsletter, (http://www.indiana.edu/~histsotl/) and there is a growing programmatic literature exploring how the discipline might respond to the challenge of SoTL (Booth, 1996; Calder, Cutler, & Kelly, 2002; Pace, 2004, 2008; Brawley, Kelly, & Timmins, 2009).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".