MétaCan
Menu
Back to cohort
Record W4229083844 · doi:10.16995/dscn.8107

Distant Approaches to the Printed Page

2022· article· en· W4229083844 on OpenAlexvenueno aff
James E. Dobson, Scott M. Sanders

Bibliographic record

VenueDigital Studies / Le champ numérique · 2022
Typearticle
Languageen
FieldArts and Humanities
TopicDigital Humanities and Scholarship
Canadian institutionsnot available
Fundersnot available
KeywordsParagraphReading (process)Object (grammar)Plot (graphics)SegmentationGestureComputer scienceArtLiteratureLinguisticsArtificial intelligencePhilosophyWorld Wide WebMathematics

Abstract

fetched live from OpenAlex

Laurence Sterne’s novel – The Life and Opinions of Tristram Shandy, Gentlemen – includes a tongue-and-cheek moment that prefigures distant reading. Near the end of the sixth volume, the narrator represents uncle Toby’s story as a meandering line with unexpected twists and predictable turns. The narrator’s precise line is inserted between two paragraphs. Its shape reminds the reader of the book’s tangential plot. But it also brings the reader back to the material contours of the story. It is a story that comes into being from the organization of lines and paragraphs on the printed page. The narrator’s precise line exists as a material object, in the middle of page 407 in volume 6 of the 1762 Lynch edition. The line gestures towards the physical space that it inhabits. In order to interpret its contours, the reader should also take into account the shape, organization and size of the printed page.This type of material analysis is under-represented in computational humanities, the majority of which has addressed segmented objects at the level of the book—actually, at the level of collections of books. The most common category of this text segmentation procedure is natural to literary scholars: the separation of individual works from within a larger collection of texts. Other categories or types of text segmentation might include the segmentation and parcellation of a longer text into its component chapters or automated algorithmically-defined procedures that ignore chapter and paragraph boundaries to cut a text or collection of texts into equally sized units of words. Segmentation enables comparison of textual objects to determine smaller effects—signals that within the larger stream of words might otherwise be lost.There has been some interest in examining individual sentences. Sarah Allison, Marissa Gemma, Ryan Heuser, Franco Moretti, Amir Tevel, and Irena Yamboliev argue that “style” exists at the level or scale of the sentence. Thematic units, however, as Mark Algee-Hewitt, Ryan Heuser, and Franco Moretti argue, might be best captured at the level of the paragraph. Sentences and paragraphs are two different units of segmentation that are both connected with linear, human reading practices. However, segmenting a text into paragraphs rids us of information about the appearance of the paragraph and its relation to the rest of the page remains occluded. Where, for example, does a particular paragraph appear in the space of the page? Are there gaps between paragraphs? Are there printed ornaments, illustrations or annotations? When digital humanists erase the footnotes from Walter Scott’s novels, the marginalia from Bunyan’s Pilgrim’s Progress and the irreverent experimental pages from Tristram Shandy, they lose the page-level context with which these texts are presented. Le roman Vie et Opinions de Tristram Shandy gentilhomme de Laurence Sterne inclut un momentironique qui préfigure la lecture à distance. Vers la fin du sixième tome, le narrateur décrit l’histoire de l’once Toby comme une ligne sinueuse avec des rebondissements inattendus et des tournures prévisibles. Les lignes précises du narrateur sont insérées entre deux paragraphes. Sa forme rappelle la tangente de l’intrigue au lecteur. Mais aussi, elle rappelle le lecteur du contour matériel de l’histoire. C’est une histoire qui voit le jour à partir d’une organisation de lignes et paragraphes sur des pages imprimées. Les lignes précises du narrateur existent en tant qu’objet matériel, dans le milieu de la page 407 du volume six de l’édition Lynch de 1762. Cette ligne fait signe à l’espace physique que celle-ci habite. Afin d’interpréter ces contours, le lecteur doit alors tenir compte de la forme, de l’organisation et de la grosseur de la page. Ce type de matériel d’analyse est sous-représenté dans le domaine des humanités informatiques, dont la majorité s’adresse à des objets segmentés au niveau du livre – et même au niveau des collections de livres. La catégorie la plus commune de cette procédure segmentée est naturelle pour les spécialistes littéraires : la séparation de travaux individuels au sein de collections de textes plus larges. D’autres catégories ou types de textes segmentés peuvent inclurent la division et le morcellement d’un texte long dans son chapitre ou des procédures algorithmiques définies et automatisées qui ignorent les limites de chapitre ou paragraphe et coupent un texte ou une collection de textes en unités de mots de la même taille. Cette division permet la comparaison d’objets textuels pour déterminer de plus petits effets – des signaux qui seraient autrement perdus dans le flux de mots plus large. Il y a de l’intérêt pour examiner les phrases individuelles. Sarah Allison, Marissa Gemma, Ryan Heuser, Franco Moretti, Amir Tevel et Irena Yamboliev soutiennent que le « style » existe au niveau ou à l’échelle de la phrase. Toutefois, comme Mark Algee-Hewitt,Ryan Heuser et Franco Moretti soutiennent, les unités thématiques pourraient être mieux capturées au niveau du paragraphe. Les phrases et paragraphes sont deux unités de segmentation différentes qui sont toutes deux connectées aux pratiques humaines de lecture linéaire. Cependant, diviser un texte en paragraphes nous enlève l’information sur l’apparence du paragraphe et les relations au reste de la page demeurent obstruées. Par exemple, où est-ce qu’un paragraphe particulier apparait sur la page? Est-ce qu’il y a des espaces entre les paragraphes? Est-ce qu’il y a des décorations, illustrations ou annotations imprimées sur la page? Lorsque les humanités numériques effacent les notes de bas de page de romans de Walter Scott, les notes marginales de Le Voyage du pèlerin de Bunyan et les pages expérimentales impertinentes de Tristam Shandy, ils perdent le contexte au niveau de la page dans lequel ces textes sont présentés. 

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesScience and technology studies, Scholarly communication
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Qualitative · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.489
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0020.000
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.182
GPT teacher head0.239
Teacher spread0.057 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designQualitative
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueDigital Studies / Le champ numériqueSame topicDigital Humanities and ScholarshipFrench-language works237,207