MétaCan
Menu
Back to cohort
Record W2060570060 · doi:10.1002/meet.14503901106

The MetaMap Project

2002· article· en· W2060570060 on OpenAlexaffabout
James M. Turner, Véronique Moal

Bibliographic record

VenueProceedings of the American Society for Information Science and Technology · 2002
Typearticle
Languageen
FieldArts and Humanities
TopicDigital and Traditional Archives Management
Canadian institutionsUniversité de Montréal
Fundersnot available
KeywordsMetadataWorld Wide WebComputer scienceData scienceKnowledge management

Abstract

fetched live from OpenAlex

For savvy Web users everywhere, there are now so many metadata initiatives that it is difficult to gain a clear understanding of what they all are and what their role is in the broader information science picture. The fact that these are usually represented by acronyms does not help, nor does the fact that although many become norms, it is not immediately clear which have this status and which do not. In addition, metadata initiatives involve many other communities than the information science community, and what the roles of each are is not always clear. Although the visual aid we are creating will necessarily be untidy and imperfect, we nevertheless hope to contribute to understanding of the relationships among metadata sets and other related information in this new and very necessary area of knowledge. The MetaMap Project is an attempt to sort out the very many efforts worldwide in working out norms and other information about metadata sets in information science. The goal is to produce a study aid that will help people interested in information science and technology to understand how the plethora of metadata initiatives spawned by the arrival of the World Wide Web are related to one another and to information studies. Sponsored by the Visual Information Research Group (GRIV) at the Université de Montréal, the project attempts to show relationships among the various norms, as well as to other pertinent information. It represents these metadata initiatives using the conventions of a subway map to help the user navigate this space, learning heavily on the conventions of the London Underground Map, noted for its clarity in helping users sort out complex reality. Each norm, metadata set, organization or other element is represented as a station on a line that has a theme. At present, lines representing themes include processes of information management: Creation, Organization, Dissemina-tion, Preservation; institutions with expertise in information management: Libraries, Archives, Museums; types of digital documentation: Text, Still Images, Moving Images, Sound. Organizations deeply involved in Web activity and metadata norms, such as the World Wide Web Consortium, OCLC, the IETF, the IEEE, and so on, are included on a separate “subway” line. Elements that have repercussions in a number of areas are represented as nodes in the “subway” network. For example, SMIL, the Synchronized Multimedia Integration Language, is common to the lines representing Text, Still Images, Moving Images, and Sound. Simpler nodes represent intersections of themes, for example where the lines representing Libraries, Archives and Museums cross the Organizations line, the nodes are respectively IFLA (International Federation of Library Associations and Institutions), ICA (International Council on Archives), and ICOM (International Council of Museums). Branch lines are created where norms such as XML or the Dublin Core spawn other norms, as a way of showing these relationships. Within a line, an attempt to order the stations to show relationships among them is made. Because of the complexity of the representation, no one criterion can be used for ordering all the lines, but those so far adopted have to do with various types of conceptual relationships among the norms such as the purpose for which they were created, their relatedness within the theme of the subway line, or chronological development (genesis) of them. Faced with the impossibility of ever being able to represent the information as clearly as we would like to within this structure, we have adopted the compromise philosophy of “well, it's much better than nothing.” To date, projected products of the initiative include a color poster which will be printed in English on one side and in French on the other. In French, the MetaMap is called MétroMéta. The other major product is a Web version of the map—actually two, one in English and one in French. SVG (Scalable Većtor Graphics) is being used to develop the Web versions, as it offers user help in navigating the information space such as zooming in and out, moving the visible area around the screen, and searching on individual acronyms. When the user passes the mouse over a station name, usually an acronym, a popup window displays the expanded name and gives other useful information such as the purpose of the metadata initiative and who is sponsoring it. Clicking on the name opens a new window and takes the user to the official Web site for the norm, metadata initiative, project, or organization. Where there is no official site, the click takes the user to the best or most complete available Web source of information. The nature of the project is such that it can never be completed. New initiatives spring up every day, it seems. Nor will the technological environment that creates the need for all this metadata activity probably ever completely settle. The inherent instability of the wonderful world of metadata means that keeping the content of our map up to date will be quite a challenge. For the moment, the best we can plan for is periodic updates, but these in turn will be dependent on the availability of human resources to carry them out. In order to be able to arrive at a product quickly, we are developing the visual representation of the MetaMap as an SVG graphic. However, what is needed in the long term is a database that will generate the map automatically from the records it contains. This will make use of exciting possibilities offered by XML and SVG as these norms themselves are developed and as we learn to use them creatively. Nevertheless, we feel that the MetaMap is a good start on organizing the information about the world we are trying to describe, and we hope it will help users in various fields of endeavor to gain an understanding of the many elements that make up that world. The Web version of the MetaMap can be found at the following site: http://mapageweb.umontreal.ca/turner/ Grateful acknowledgement is made to CoRIMedia, a consortium for research in image, video and multimedia indexing and navigation, which is financed by Valorisation-Recherche Québec, an initiative of the government of Québec, and industrial partners. CoRIMedia in turn is financing this project.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.013
metaresearch head score (Gemma)0.026
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.077
Threshold uncertainty score0.258

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0130.026
Meta-epidemiology (narrow)0.0030.002
Meta-epidemiology (broad)0.0020.005
Bibliometrics0.0080.009
Science and technology studies0.0030.003
Scholarly communication0.0150.018
Open science0.0080.018
Research integrity0.0050.006
Insufficient payload (model declined to judge)0.0770.082

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.025
GPT teacher head0.225
Teacher spread0.200 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2002
Admission routes2
Has abstractyes

Explore more

Same venueProceedings of the American Society for Information Science and TechnologySame topicDigital and Traditional Archives ManagementFrench-language works237,207