Content-Based Art Retrieval (C-BAR)
Bibliographic record
Abstract
Andersen & Erikstad -Making a national census coding system internationally comparable 23 Berezhnoy, Postma & Van den Herik -Computerized visual analysis of paintings 28 Bergboer, Postma & Van den Herik -Visual object detection for the cultural heritage 33 Berger -Microhistory and quantitative data analysis 39 Boot -Advancing digital scholarship using EDITOR 43 Boughida -CDWA lite for Cataloguing Cultural Objects (CCO): A new XML schema for the cultural heritage community 49 Breure -PROGENETOR: An editorial framework for reuse of XML content 57 Broadway -The early letters of The Royal Society 1657-1741: Managing diversity and complexity 64 Van den Broek, Kok, Hoenkamp, Schouten, Petieta & Vuurpijl -Content-Based Art Retrieval (C-BAR) 70 Van den Broek, Wiering & Van Zwol -Backing the Right Horse: Benchmarking XML editors for text-encoding 78 Brunnhofer & Kropač -Digital archives in a virtual world 83 Burkard -Collaboration on medieval charters -Wikipedia in the humanities?91 Burrows -Reinventing the humanities in a networked environment: the Australian Network for Early European Research 95 Clausen -Digitising parish registers -principles and methods 100 Delve & Healey -Is there a role for data warehousing technology in historical research?106 Doorenbosch -Computer science and the Dutch cultural heritage 112 Fogelvik -Large longitudinal, nominative databases in historical research 119 Garskova -Towards a standard for MA programs in historical computing: (the experience of Russian and CIS universities) 123 Glavatskaya -Indigenous peoples of the North-western Siberia: Ethnohistorical mapping 126 Gregory -Creating analytic results from historical GIS 131 Gruber -Occupational migration in Albania in the beginning of the 20th century 136 Heller & Vogeler -Modern information retrieval technology for historical documents 143 Hoekstra -Integrating structured and unstructured searching in historical sources 149 Ivanovs & Varfolomeyev -Editing and exploratory analysis of medieval documents by means of XML technologies 155 De Jong, Rode & Hiemstra -Temporal language models for the disclosure of historical text 161 Juola -Language change and historical inquiry 169 Kröll -Not ready for the Semantic Web: A fi eld study of subject gateways on Contemporary History 176 Laloli -Moving through the city: residential mobility and social segregation in Amsterdam 1890-1940 182 Lopes-Historical geographic data dissemination through the web: the site Atlas and future developments towards its interoperability 190 Melms -Reconstructing lost spaces.Affordably, that is 194 Mirzaee, Iverson, Hamidzadeh -Computational representation of semantics in historical documents 199 Nagypál -History ontology building: The technical view 207 Ordelman, De Jong, Huijbregts & Van Leeuwen -Robust audio indexing for Dutch spokenword collections 215 Pasqualis Dell Antonio -From the roman eagle to E.A.G.L.E.: harvesting the web for ancient epigraphy 224 Perstling -Layers and dimensions.The representation of complex structured sources 229 Petty -Transnational histories in Roshini Kempadoo's ghosting.Cyber)Race identities.237 Pieken -Jewish life in Germany from 1914 to 2004 -The story of the Chotzen family 243 Robichaud -The old Montréal heritage inventory database: Toward a renewed collective memory 246 Tschauner & Siveroni Salinas -On the ground and '6 feet under'.Mobile GIS and photogrammetric approaches to building 3D archaeological spatial databases in the fi eld 250 Valetov -World museums on the Internet: A brief overview 255 Verheusen -National digital repository for cultural heritage institutions 263 Voegler -Virtual libraries and thematic gateways in German history: Strategies and perspectives 267 Weller -A new approach: The arrival of informational history 273 Wiering, Crawford & Lewis -Creating an XML vocabulary for encoding lute music 279 Wouters -Writing history in the virtual knowledge studio for the humanities and social sciences 288 Zandhuis -Towards a genealogical ontology for the Semantic Web 296 Zeldenrust -DIMITO: Digitization of rural microtoponyms at the Meertens Instituut 301 Patricia AlkhovenMany Cultural Heritage Institutions are starting to digitize parts of their collections.They have often very few staff available for these 'new' activities.Often without any training or preparation they start scanning objects.Very little attention is payed to standards, interoperability, workfl ow effi ciency, cooperation with other institutes or even the quality of the product in terms of scholarly content, image quality, usability etc.This paper explanes why training is necessary before huge investments in time and money are made and end-products appear to be disappointing.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.013 | 0.011 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.003 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.003 | 0.001 |
| Insufficient payload (model declined to judge) | 0.080 | 0.082 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".