MétaCan
Menu
← Back to cohort
Record W6949619847 · doi:10.5281/zenodo.14277812

Insect DNA Barcode and Image Dataset

2024· dataset· en· W6949619847 on OpenAlexaboutno aff

Bibliographic record

VenueZenodo (CERN European Organization for Nuclear Research) · 2024
Typedataset
Languageen
FieldMedicine
TopicPrenatal Screening and Diagnostics
Canadian institutionsnot available
Fundersnot available
KeywordsBarcodeDNA barcodingPattern recognition (psychology)DNAVector (molecular biology)DNA sequencingImage (mathematics)Genus

Abstract

fetched live from OpenAlex

The data utilized in our experiments was obtained from the Barcode of Life Data System (BOLD), which is a cloud-based data storage and analysis platform developed at the Centre for Biodiversity Genomics in Canada. The insect_dataset.mat consists of 32424 image samples of insect species from four Insecta orders, Diptera, Coleoptera, Lepidoptera and Hymenoptera, each associated with a DNA barcode sequence of that sample. The unseen_insect_dataset.mat consists of 40050 image samples of insects from the same order, but all don't have an indicated species in the BOLD System, so they are real unclassified species (at the time of the dataset creation), having only the genus available, each one is also associated with a DNA barcode sequence of that sample. Description of the .mat: # Insect Dataset * all_images: vector containing the 32424 64x64x3 images (RGB) pre normalized of the insects* all_dnas: vector containing the 32424 DNA barcodes in one-hot encoding 658x5* all_labels: vector containing the species label for the corresponding DNA and image* all_boldids: vector of strings containing the id from boldsystemsv3 (https://v3.boldsystems.org/) they can be used to download from boldsystems the original DNA barcodes and the full size images and other data related to the sample* train_loc: indices of the training samples in all_dnas, all_images, all_labels, all_boldids* val_seen_loc: indices of the validation samples in all_dnas, all_images, all_labels, all_boldids that contain described(seen) species* val_unseen_loc: indices of the validation samples in all_dnas, all_images, all_labels, all_boldids that contain undescribed(unseen) species* test_seen_loc: indices of the test samples in all_dnas, all_images, all_labels, all_boldids that contain described(seen) species* test_unseen_loc: indices of the test samples in all_dnas, all_images, all_labels, all_boldids that contain undescribed(unseen) species* species2genus: the vector contains at index i the genus label of species with label i (e.g. species i has genus species2genus[i])* described_species_labels_train: vector containing the labels of species that appear in the training set* described_species_labels_trainval: vector containing the labels of species that appear in the training set and/or the validation set* all_dna_features_cnn_original: vector of features extractedfrom DNA nucleotides with the method of Badirli, S., Picard, C. J., Mohler, G.,Richert, F., Akata, Z., & Dundar, M. (2023). Classifying the unknown: Insect identification with deep hierarchical Bayesian learning. Methods in Ecology and Evolution, 14,1515-1530. https://doi.org/10.1111/2041-210X.14104* all_image_features_resnet: vector of features extracted from the insect images with the method of the same paper as the all_dna_features_cnn_original with a pretrained resnet101* all_dna_features_cnn_new: vector of features extracted from DNA nucleotides with our CNN* all_image_features_gan: vector of features extracted from the insect images with out method using a ReACGAN Description of the .mat: # Unseen Insect Dataset * all_images: vector containing the 40050 64x64x3 images (RGB) pre normalized of the insects* all_dnas: vector containing the 40050 DNA barcodes in one-hot encoding 658x5 * all_string_dnas: vector containing the 40050 DNA barcodes in string format* all_genus_labels: vector containing the species label for the corresponding DNA and image* all_boldids: vector of strings containing the id from boldsystemsv3 (https://v3.boldsystems.org/) they can be used to download from boldsystems the original DNA barcodes and the full size images and other data related to the sample* all_dna_features_cnn_original: vector of features extractedfrom DNA nucleotides with the method of Badirli, S., Picard, C. J., Mohler, G.,Richert, F., Akata, Z., & Dundar, M. (2023). Classifying the unknown: Insect identification with deep hierarchical Bayesian learning. Methods in Ecology and Evolution, 14,1515-1530. https://doi.org/10.1111/2041-210X.14104* all_image_features_resnet: vector of features extracted from the insect images with the method of the same paper as the all_dna_features_cnn_original with a pretrained resnet101* all_dna_features_cnn_new: vector of features extracted from DNA nucleotides with our CNN* all_image_features_gan: vector of features extracted from the insect images with out method using a ReACGAN Note: all arrays and locs are 1-indexed like in MATLAB. Note: the features were extracted with the same model for both the insect dataset and the unseen insect dataset.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.003
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.028
Threshold uncertainty score0.094

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.003
Meta-epidemiology (narrow)0.0020.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.003
Science and technology studies0.0010.000
Scholarly communication0.0010.001
Open science0.0020.002
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0280.050

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.039
GPT teacher head0.280
Teacher spread0.241 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueZenodo (CERN European Organization for Nuclear Research)→Same topicPrenatal Screening and Diagnostics→French-language works237,207→