Bibliographic record
Abstract
Thomas Hughes's Tom Brown's Schooldays, ChatGPT, and Academic Integrity Tom Ue (bio) Chatgpt is all the rage. In the Chronicle of Higher Education (12 May 2023), Owen Kichizo Terry, an undergraduate student at Columbia University, observes how easy it is "to use AI to do the lion's share of the thinking while still submitting work that looks like your own." Terry advocates for a "massive structural change" in colleges for them "to keep training students to think critically." By this, Terry refers to actively embracing AI's role in the writing process or creating "a split between assignments on which using AI is encouraged and assignments on which using AI can't possibly help." Many a humanist and scientist share Terry's conviction that AI is here to stay, and they have ruminated thoughtfully on its affordances and limitations. In an editorial for our discipline's pre-eminent journal, PMLA, for example, Wai Chee Dimock imagines our co-existence with AI. "Literature from Gilgamesh on," she says, "has taught us about the human assault on the nonhuman world. It has also taught us the art of assisted survival by making kin with nonhuman beings. The emergence of AI at this moment of crisis makes that art all the more urgent" (453). Lauren M.E. Goodlad responds to Dimock, also in PMLA, by describing AI's major drawback, "impressive strides" notwithstanding: "Lacking sentience, emotion, common sense, imagination, and a model of the world, these powerful pattern-finders cannot cognize the data points they extrapolate" (317).1 Meanwhile, in the top science journal, Nature, Eva A.M. Van Dis et al. share Goodlad's reservations, and they similarly suggest the need to put AI in its place: "The focus should be on embracing the opportunity and managing the risks. We are confident that science will find a way to benefit from conversational AI without losing the many important aspects that render scientific work one of the most profound and gratifying enterprises: curiosity, imagination and discovery" (226). [End Page 21] Conversations about academic integrity are not new. Almost two centuries before ChatGPT, the eponymous character of Thomas Hughes's Tom Brown's Schooldays (1857) had vulgus-books. One evening, Tom and his schoolmates Arthur and Martin are beavering away at their vulgus task, "a short exercise, in Greek or Latin verse, on a given subject, the minimum number of lines being fixed for each form" (259). Here's how the assignment works: The master of the form gave out at fourth lesson on the previous day the subject for next morning's vulgus, and at first lesson each boy had to bring his vulgus ready to be looked over; and with the vulgus, a certain number of lines from one of the Latin or Greek poets then being construed in the form had to be got by heart. The master at first lesson called up each boy in the form in order, and put him on in the lines. If he couldn't say them, or seem to say them, by reading them off the master's or some other boy's book who stood near, he was sent back, and went below all the boys who did so say or seem to say them; but in either case his vulgus was looked over by the master, who gave and entered in his book, to the credit or discredit of the boy, so many marks as the composition merited. (259–60; my emphases) Hughes's narrator gestures, through the incantatory use of the expression "seem to say them," at the pupils' inaudible murmurings, and, through their "reading them off the master's or some other boy's book," at their wandering eyes. Cheating, Hughes goes on to suggest, is by no means confined to oral assessments, and he describes four methods by which students complete their tasks. That a master is tasked to come up with 114 subjects a year (i.e., three a week for each of the school year's thirty-eight weeks) ensures that some of them would be repeated. Pupils "meet and rebuke this bad habit" by handing down their exercises so that "the popular...
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.027 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.016 | 0.008 |
| Scholarly communication | 0.012 | 0.007 |
| Open science | 0.002 | 0.008 |
| Research integrity | 0.006 | 0.015 |
| Insufficient payload (model declined to judge) | 0.070 | 0.025 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".