MétaCan
Menu
Retour à la cohorte
Enregistrement W6968129748 · doi:10.5281/zenodo.14604271

TEI Technical Infrastructure

2018· article· en· W6968129748 sur OpenAlexaff

Notice bibliographique

RevueZenodo (CERN European Organization for Nuclear Research) · 2018
Typearticle
Langueen
DomaineComputer Science
ThématiqueMathematics, Computing, and Information Processing
Établissements canadiensUniversity of Victoria
Organismes subventionnairesnon disponible
Mots-clésXSLTXMLJSONXQueryInteroperabilityScripting languageMetadata

Résumé

récupéré en direct d'OpenAlex

The main goal of the Text Encoding Initiative (TEI) is the development of a standard for encoding of textual phenomena and manifests itself in the TEI Guidelines and the derived (formal) schemata. For the production and maintenance of these Guidelines and its various file formats, as well as for processing any kind of TEI document, a larger ecosystem of tools, methods, and technical service infrastructure have evolved around the TEI standard. Since the most common serialization format for TEI documents is XML, a lot of generic XML tools are employed, e.g., Apache ANT as a build tool, Saxon as XML processor, and XSLT and XQuery as scripting and query languages. In fact, the most prominent piece of software that the TEI Consortium produces are the TEI Stylesheets, a set of XSL stylesheets that convert to and from TEI files. Supported import and export formats include docx, markdown, HTML, PDF, and many more. But the TEI Stylesheets not only convert ‘regular’ TEI documents but also TEI ODD customization files, acting as an ODD processor to produce both documentation and schemata from an ODD customization—again in various possible output formats. The TEI stylesheets are at the core of most available TEI transformation services including the OxGarage RESTful web service, the oXygen TEI frameworks, the MEI Customization Service, and the TEI Jenkins continuous integration servers which test the TEI build process. Other services aim at providing an online editor for TEI ODD customizations rely on the aforementioned OxGarage services to handle these transformations e.g., produce JSON serializations of the TEI specifications for their internal data structures, and return schemata and documentation to the user. The whole multitude of tools and services can best be understood by looking at the TEI Git repositories that are hosted at GitHub under https://github.com/TEIC. The most prominent are the TEI Guidelines and Stylesheets themselves, while CETEIcean is another rising star. CETEIcean is a pure CSS and Javascript renderer for TEI files which facilitates not only the display, but also the online editing of TEI documents. It is a smart front end library which can be used in one’s own project by simply including the needed javascript and CSS files. Most of the aforementioned tools depend on other services and software. It is the aim of the poster to illustrate these dependencies and to give a thorough overview of the tools and services (actively) maintained by the TEI Consortium. Such an overview would help to distinguish the various tools and services for those members of the TEI community, who find it hard to know what the relationships are between Roma, OxGarage, the Stylesheets, and the Guidelines. While the more tech savvy members would hopefully appreciate these insights for installing and running these services on their own hardware. Finally, the TEI Consortium itself would benefit from better documentation of their service architecture and the feedback of their users gained by this poster.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,006
score de la tête « metaresearch » (Gemma)0,013
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesCommunication savante
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Méthodes · Signal consensuel: aucune
Score de désaccord entre enseignants0,992
Score d'incertitude au seuil0,961

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0060,013
Méta-épidémiologie (sens strict)0,0020,001
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0070,007
Études des sciences et des technologies0,0030,001
Communication savante0,0080,007
Science ouverte0,0060,008
Intégrité de la recherche0,0030,004
Charge utile insuffisante (le modèle a refusé de juger)0,2870,372

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,021
Tête enseignante GPT0,243
Écart entre enseignants0,222 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeSans objet
Domainenon disponible
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2018
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueZenodo (CERN European Organization for Nuclear Research)Même sujetMathematics, Computing, and Information ProcessingTravaux en français237 207