MétaCan
Menu
← Retour à la cohorte
Enregistrement W6911800849 · doi:10.5281/zenodo.13255763

HTR to TEI: Strategies for Transformation, Up-conversion, and Data Enrichment of Transcripts

2024· article· en· W6911800849 sur OpenAlexaboutno aff

Notice bibliographique

RevueZenodo (CERN European Organization for Nuclear Research) · 2024
Typearticle
Langueen
DomaineArts and Humanities
ThématiqueDigital Humanities and Scholarship
Établissements canadiensnon disponible
Organismes subventionnairesUK Research and Innovation
Mots-clésWorkflowProcess (computing)XMLPresentation (obstetrics)Transformation (genetics)Point (geometry)Data transformationData processing

Résumé

récupéré en direct d'OpenAlex

The ideas for this paper proposal originate partly from the Evolving Hands project, funded by the AHRC-NEH New Directions in Digital Scholarship for Cultural Institutions programme, and partly from combined decades of data processing experience by the authors. The topic is not the project itself, or the process of Handwritten Text Recognition (HTR), but the transformation, up-conversion, and enrichment of data produced by these processes. This presentation will be of benefit to those considering using HTR where further processing of the transcripts will be needed. Introduction and Background Due to the time constraints of a short presentation, the authors, will only spend a limited amount of time on introducing the project and background issues. We will not explain how HTR works or recent assessments of its end-to-end performance in pattern recognition as these exist elsewhere. Instead, we will start at the point of export, noting that most HTR systems, and especially the READ Co-Op’s Transkribus platform that the project selected, use the Page XML for export of HTR transcripts. The Evolving Hands project looked specifically at the export of data from Transkribus and the workflows for transformation of these transcripts to TEI P5 XML, the up-conversion of them by adding additional markup, annotation, and contextual information. Transformation There are many approaches to this form of data processing and conversion, and the transformation of data from the outputs of HTR conversion. Transkribus exports to Page and Alto XML formats, and also has a conversion to TEI. Part of the work done on this project examined various routes for export and conversion to TEI, noting that at the time, the direct conversion had a number of limitations. Hence, some of the case studies in the project used the more complete Page XML export and then used a more recent version of Dario Kampkaspar’s Page2TEI XSLT conversion for transforming the HTR transcripts. The paper will look at some of the options in transformation of files to TEI P5 XML and will also investigate other considerations of transformation for potential users. XML Up-Conversion Up-conversion when talking about XML data formats (and not video) is usually understood as “the generation of a format with detailed markup from a format with less-detailed or no markup, where it is necessary to generate the additional markup by recognizing structural patterns that are implicit in the textual content itself”. Usually this involves the transformation of data from one format often (but not necessarily) to the same format. During the process the data is supplemented with additional data provided through methods such as: the probabilistic intuiting of existing data structures based on less-detailed markup, the wholesale replacement of existing hierarchies with more detailed and expressive substitutes, or the determination of extra annotation based on retrieving data from external sources using contextual clues. While any programming language can suffice for XML up-conversion, XSLT has often been favoured. While all have their different strengths, such conversions can be completed in any language, as long as the conversions are transparent, complete, and reproducible. The paper will look at the XML up-conversion of a number of data sets to circumvent limitations in HTR processes. One example uses HTR for the conversion of structured print volumes of the Records of Early English Drama (REED) project which are unsuitable for OCR because of the intra-word formatting used to indicate the expansion of abbreviations in late-medieval and Early Modern edited records concerning performance history. However, HTR can use a cypher-based system in which those characters that are supplied in italics are recognised by the HTR process as a completely different character. We used non-anglophone currency symbols and similar characters unlikely to appear in pre-modern manuscripts and thus not the REED collections. The recognised cipher characters can be reverted into markup versions of the correct characters. The results can be merged together with the markup of the whole word up-converted into a more detailed format (using tei:choice). The result of this workflow is the preservation of the intellectual content of the original editor for the expansion of any pre-modern abbreviation in the collection. A separate use is for an XSLT Stylesheet which converts the abbreviations and expansions of REED collections into an abbreviation wordlist for other purposes. Data Enrichment The third part of the paper will look at some of the semi-automatic forms of data enrichment investigated by the project using tools created by other projects where the authors are directly involved. For example, the use of LEAF-Writer whether via the LEAF project’s LEAF-Writer Commons or embedded in the LEAF-VRE system itself to provide Named Entity Recognition (NER) as a form of data enrichment. The paper will explain the underlying process which is using the Named Entity Relationship and Vetting Environment (NERVE) created by the LINCS project. The presentation will highlight the ease of use of such tools for data enrichment, especially where LEAF-Writer has a dedicated import for HTR-generated TEI P5 XML files. Works Cited Cummings, J., Jakacki, D., et al. (2023) ‘Building Workflows for HTR to TEI Up-Conversion and Enhancement’, TEI MEC Joint Conference 2023. Haupt, S., Stührenberg, M. (2010) “Automatic upconversion using XSLT 2.0 and XProc: A real world example.” Presented at Balisage: The Markup Conference 2010, Montréal, Canada, August 3 - 6, 2010. In Proceedings of Balisage: The Markup Conference 2010. Balisage Series on Markup Technologies, vol. 5. https://doi.org/10.4242/BalisageVol5.Haupt01. Jakacki, D., Brown, S., Cummings, J., Ilovan, M., (2023): ‘LEAF: Developing Streamlined Digital Scholarly Workflows with the Linked Editing Academic Framework’ workshop at the Digital Humanities 2023 Annual Meeting. Kay, M. (2004) "Up-conversion using XSLT 2.0.", XML 2004 Conference http://www.saxonica.com/papers/ideadb-1.1/mhk-paper.xml. Vidal, E., Toselli, A., Ríos-Vila, A., Calvo-Zaragoza, J. (2023) ‘End-to-End page-Level assessment of handwritten text recognition’ , Pattern Recognition, Volume 142, October 2023, 109695. https://doi.org/10.1016/j.patcog.2023.109695

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,017
score de la tête « metaresearch » (Gemma)0,055
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: aucune
GenreSignal candidat: Méthodes · Signal consensuel: Méthodes
Score de désaccord entre enseignants0,028
Score d'incertitude au seuil0,093

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0170,055
Méta-épidémiologie (sens strict)0,0020,002
Méta-épidémiologie (sens large)0,0010,002
Bibliométrie0,0050,006
Études des sciences et des technologies0,0020,004
Communication savante0,0110,011
Science ouverte0,0030,011
Intégrité de la recherche0,0020,005
Charge utile insuffisante (le modèle a refusé de juger)0,0280,029

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,109
Tête enseignante GPT0,274
Écart entre enseignants0,164 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2024
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueZenodo (CERN European Organization for Nuclear Research)→Même sujetDigital Humanities and Scholarship→Travaux en français237 207→