‘The Apex of Hipster XML GeekDOM’
Bibliographic record
Abstract
If the notion of the methodological commons is as centrally located as we believe it to be in any visualization accurately depicting the intellectual structure of the digital humanities and digital literary studies (McCarty 2005, 119), then so, too, must be the community itself whose members provide that which populates the commons. As an interdiscipline, humanities computing has always well-understood its methodologies; indeed, the digital humanities (of which digital literary studies is a part), more generally, have made a virtue of the way in which they render explicit and tangible the theoretical models that govern the representative and analytical endeavour of their fields via computational application. So, too, have those in the field understood and documented its formal structures and institutional manifestations, a chief example being the Text Encoding Initiative itself. Less explicitly rendered and less formally documented–though intuited by its chief practitioners and builders–is the exact nature of the community itself, its depth and breadth, its own centre and, perhaps more important in a field whose embrace of interdisciplinarity is far from self-serving, its periphery and those aspects of which promise to become central. This article presents work carried out in conjunction with the Text Encoding Initiative Consortium, a foundation of many digital literary studies projects, work that seeks to document the full nature of its community, from the institutional and research project groups that comprise the formal consortium at centre to those who appear on the other side of the easily-permeable periphery that separates it from the centre, largely individual practitioners in areas hitherto not closely identified with the digital humanities but clearly sharing methods and tools, thus suggesting their place in the same communities of practice, as they are members of the same methodological commons. This methodological approach is drawn from marketing and organizational behavior, manifest in social networking, in the study of viral marketing campaigns conducted in online environments. The method for this work was centred around a viral marketing experiment designed to showcase the TEI and novel ways that it can be used to encode different kinds of text. At the heart of the experiment was a Bob Dylan song and its associated video which incorporated text; encoded text was overlaid and the video was posted to YouTube and a blog with links to the TEI website with analysis of traffic patterns carried out.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.030 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.013 | 0.023 |
| Open science | 0.004 | 0.011 |
| Research integrity | 0.003 | 0.005 |
| Insufficient payload (model declined to judge) | 0.062 | 0.032 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".