HTR to TEI: Strategies for Transformation, Up-conversion, and Data Enrichment of Transcripts
Bibliographic record
Abstract
The ideas for this paper proposal originate partly from the Evolving Hands project, funded by the AHRC-NEH New Directions in Digital Scholarship for Cultural Institutions programme, and partly from combined decades of data processing experience by the authors. The topic is not the project itself, or the process of Handwritten Text Recognition (HTR), but the transformation, up-conversion, and enrichment of data produced by these processes. This presentation will be of benefit to those considering using HTR where further processing of the transcripts will be needed. Introduction and Background Due to the time constraints of a short presentation, the authors, will only spend a limited amount of time on introducing the project and background issues. We will not explain how HTR works or recent assessments of its end-to-end performance in pattern recognition as these exist elsewhere. Instead, we will start at the point of export, noting that most HTR systems, and especially the READ Co-Op’s Transkribus platform that the project selected, use the Page XML for export of HTR transcripts. The Evolving Hands project looked specifically at the export of data from Transkribus and the workflows for transformation of these transcripts to TEI P5 XML, the up-conversion of them by adding additional markup, annotation, and contextual information. Transformation There are many approaches to this form of data processing and conversion, and the transformation of data from the outputs of HTR conversion. Transkribus exports to Page and Alto XML formats, and also has a conversion to TEI. Part of the work done on this project examined various routes for export and conversion to TEI, noting that at the time, the direct conversion had a number of limitations. Hence, some of the case studies in the project used the more complete Page XML export and then used a more recent version of Dario Kampkaspar’s Page2TEI XSLT conversion for transforming the HTR transcripts. The paper will look at some of the options in transformation of files to TEI P5 XML and will also investigate other considerations of transformation for potential users. XML Up-Conversion Up-conversion when talking about XML data formats (and not video) is usually understood as “the generation of a format with detailed markup from a format with less-detailed or no markup, where it is necessary to generate the additional markup by recognizing structural patterns that are implicit in the textual content itself”. Usually this involves the transformation of data from one format often (but not necessarily) to the same format. During the process the data is supplemented with additional data provided through methods such as: the probabilistic intuiting of existing data structures based on less-detailed markup, the wholesale replacement of existing hierarchies with more detailed and expressive substitutes, or the determination of extra annotation based on retrieving data from external sources using contextual clues. While any programming language can suffice for XML up-conversion, XSLT has often been favoured. While all have their different strengths, such conversions can be completed in any language, as long as the conversions are transparent, complete, and reproducible. The paper will look at the XML up-conversion of a number of data sets to circumvent limitations in HTR processes. One example uses HTR for the conversion of structured print volumes of the Records of Early English Drama (REED) project which are unsuitable for OCR because of the intra-word formatting used to indicate the expansion of abbreviations in late-medieval and Early Modern edited records concerning performance history. However, HTR can use a cypher-based system in which those characters that are supplied in italics are recognised by the HTR process as a completely different character. We used non-anglophone currency symbols and similar characters unlikely to appear in pre-modern manuscripts and thus not the REED collections. The recognised cipher characters can be reverted into markup versions of the correct characters. The results can be merged together with the markup of the whole word up-converted into a more detailed format (using tei:choice). The result of this workflow is the preservation of the intellectual content of the original editor for the expansion of any pre-modern abbreviation in the collection. A separate use is for an XSLT Stylesheet which converts the abbreviations and expansions of REED collections into an abbreviation wordlist for other purposes. Data Enrichment The third part of the paper will look at some of the semi-automatic forms of data enrichment investigated by the project using tools created by other projects where the authors are directly involved. For example, the use of LEAF-Writer whether via the LEAF project’s LEAF-Writer Commons or embedded in the LEAF-VRE system itself to provide Named Entity Recognition (NER) as a form of data enrichment. The paper will explain the underlying process which is using the Named Entity Relationship and Vetting Environment (NERVE) created by the LINCS project. The presentation will highlight the ease of use of such tools for data enrichment, especially where LEAF-Writer has a dedicated import for HTR-generated TEI P5 XML files. Works Cited Cummings, J., Jakacki, D., et al. (2023) ‘Building Workflows for HTR to TEI Up-Conversion and Enhancement’, TEI MEC Joint Conference 2023. Haupt, S., Stührenberg, M. (2010) “Automatic upconversion using XSLT 2.0 and XProc: A real world example.” Presented at Balisage: The Markup Conference 2010, Montréal, Canada, August 3 - 6, 2010. In Proceedings of Balisage: The Markup Conference 2010. Balisage Series on Markup Technologies, vol. 5. https://doi.org/10.4242/BalisageVol5.Haupt01. Jakacki, D., Brown, S., Cummings, J., Ilovan, M., (2023): ‘LEAF: Developing Streamlined Digital Scholarly Workflows with the Linked Editing Academic Framework’ workshop at the Digital Humanities 2023 Annual Meeting. Kay, M. (2004) "Up-conversion using XSLT 2.0.", XML 2004 Conference http://www.saxonica.com/papers/ideadb-1.1/mhk-paper.xml. Vidal, E., Toselli, A., Ríos-Vila, A., Calvo-Zaragoza, J. (2023) ‘End-to-End page-Level assessment of handwritten text recognition’ , Pattern Recognition, Volume 142, October 2023, 109695. https://doi.org/10.1016/j.patcog.2023.109695
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.055 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.005 | 0.006 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.011 | 0.011 |
| Open science | 0.003 | 0.011 |
| Research integrity | 0.002 | 0.005 |
| Insufficient payload (model declined to judge) | 0.028 | 0.029 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".