Building Next-Generation Collections: Natural History Specimens, Just One Click Away!
Bibliographic record
Abstract
Digitisation has made significant advances in many natural history collections since the 1980s. The Vertebrate Zoology Collections team of the Canadian Museum of Nature (CMNVZC; ca. 1,250,000 catalogued specimens) has the ambition to go fully digital with our physical objects and associated data. Organising CMNVZC data electronically (primary digitisation) through computerisation for collection management purposes was initiated in 1972 and systematically implemented since the 1980s. This databasing process involved several stages, each with its own objectives and challenges. It resulted in ca. 100% of the CMNVZC being now digitised and core specimen data being retrievable from the Web (e.g., GBIF, and VertNet). Digitising requires regular updates to reflect the changing needs of the collections-based research community, and to capitalise on new opportunities that arise with the advances in technology. In this digital age, improving collections accessibility and usability through realistic and sustainable digitisation, while avoiding the downside of information overload, remains the most pressing challenge. Increasing CMNVZC accessibility necessitates further consolidation and information standardisation of various types (e.g. collecting data) to be retrieved from several sources (e.g., field notes, original data sheets, and maps). Optimising collections usability can be achieved by adding value to existing records (secondary digitisation) by means of additional information as mentioned above, georeferencing, as well as 2D and 3D imaging. Virtual sharing of 3D specimen images allows for remote examination of specimens usually inaccessible through loans, such as type and rare specimens, and the possibility for morphometric analyses. Digital imaging of the vertebrate collection, however, represents a major challenge given the complexity and variation of shapes and sizes among specimens. Limitations of current 3D surface imaging technology, none of which have been specifically designed for natural history specimens, hamper CMNVZC imaging workflows. Digital tools are key to the success of increasing usability of natural history collections and play an important role in preserving information. Digitisation activities should endeavour to improve online access of physical objects and their full array of data with optimized usability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.004 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.013 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".