MétaCan
Menu
Back to cohort
Record W2138920268

Digitalizzazione e linguaggi di marcatura

2012· article· it· W2138920268 on OpenAlexaboutno aff
Enrico Seta

Bibliographic record

VenueBollettino AIB (1992-2012) · 2012
Typearticle
Languageit
FieldComputer Science
TopicLibrary Science and Information Systems
Canadian institutionsnot available
Fundersnot available
KeywordsDigitizationWorld Wide WebComputer scienceSGMLXMLStyle sheetInformation technologyKnowledge managementData scienceDocument Structure Description
DOInot available

Abstract

fetched live from OpenAlex

The Digital Library is - above all - a very favourable opportunity to enrich, in different ways, the knowledge and professional skills: librarians should not miss this chance. In fact, the digitization projects ask for project management skills, awareness of new technologies and related costs, interaction with professionals belonging to different fields, and, last but not least, ability to attract adequate funds. Furthemore, the partnership with some dynamic companies belonging to the information technology area can deeply change the working style and practices in the libraries. On the descriptive ground, the article does not survey some momentous subjects of digital libraries (e.g. copyright and digital resources preservation) because it focuses only on a specific topic: the comparative analysis between technologies, costs and outcomes of capture of digital images and capture (or creation) of digital texts. When the collection is made up of textual documents rich of informative contents, it should be endeavoured to get an electronic text, because it's a good investment in the long term. At present, satisfactory results are achievable through ICR and fuzzy searching software. But the final aim of a digitization project, for such collections, should be to get a structured text, and the structure should be a logical and permanent one. The article explains the main features of SGML standard and the more recent advances, especially the innovative attributes of XML. This new metalanguage, born in the middle of Web explosion, drag into the Web the SGML ability to transport information. This feature should arise the interest of information professionals and librarians. XML electronic documents can suit very well preservation and access requirements. Furthemore, they hold a cardinal position in the Web of the future: in this forthcoming environment, structured texts, jontly with metadata, will play a key-role in the exploitation of searching functions. HTML and PDF are probably bound to fall into decline in the semantic Web. This tendencies can be surveyed into the documents of the web standards community , especially the W3C metadata activity. The author refers about experiences and research in the field of SGML/XML application to digitization projects. In particular, he refers about the Electronic Text Centers activity, underlining the role of library community within an interdisciplinary model. In some countries like US, UK, and Canada text encoding already is a part of librarians skills. But it is important to remark that this is a natural evolution (and enrichment) of cataloging skills and methodology. What should be required is the accurate knowledge of the markup languages and the ability to write DTDs (Document Type Definitions), not only a generic awareness of the subject. The ability to typify is quite similar to cataloguing skills, and a good typification cannot lack the experience of the information intermediaries. The computer programmers need the interaction with information specialists to capture the logical structure of the text and therefore to give a better design to the whole information system. XML, together with metadata, represents the ground of an interdisciplinary activity, and a chance for librarians to play a dynamic role into the digital environment.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.004
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesScholarly communication
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.991
Threshold uncertainty score0.063

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.004
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0040.005
Science and technology studies0.0030.003
Scholarly communication0.0090.003
Open science0.0010.003
Research integrity0.0010.004
Insufficient payload (model declined to judge)0.0190.009

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.021
GPT teacher head0.237
Teacher spread0.216 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designTheoretical or conceptual
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2012
Admission routes1
Has abstractyes

Explore more

Same venueBollettino AIB (1992-2012)Same topicLibrary Science and Information SystemsFrench-language works237,207