MétaCan
Menu
Retour à la cohorte
Enregistrement W4286239843 · doi:10.5281/zenodo.6866942

Datasets for Data-Centric Classification and Clustering

2022· article· en· W4286239843 sur OpenAlexaboutno aff
Lars Schmarje, Monty Santarossa, Simon‐Martin Schröder, Claudius Zelenka, Rainer Kiko, Jenny Stracke, Nina Volkmann, Reinhard Koch

Notice bibliographique

RevueZenodo (CERN European Organization for Nuclear Research) · 2022
Typearticle
Langueen
DomaineComputer Science
ThématiqueAdvanced Clustering Algorithms Research
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésCluster analysisComputer scienceData miningData scienceArtificial intelligence

Résumé

récupéré en direct d'OpenAlex

This is the technical descriptions of the used datasets in the paper "A data-centric approach for improving ambiguous labels with combined semi-supervised classification and clustering" (https://arxiv.org/abs/2106.16209). The source code is available at https://github.com/Emprime/dc3 . We provide as summary taken from the original work and technical descriptions for all datasets: The Plankton dataset was introduced in "Fuzzy Overclustering: Semi-supervised classification of fuzzy labels with overclustering and inverse cross-entropy" (https://doi.org/10.3390/s21196661). The dataset contains 10 plankton classes and has multiple labels per image due to the help of citizen scientists. In contrast to the previous work, we include fuzzy images in the training and validation set and do not enforce a class balance which results in a slighlty different data split. Moreover, we preprocessed the data by recentering the images and removing artifacts like scale bars. The Turkey dataset was used in "Learn to train: Improving training data for a neural net- work to detect pecking injuries in turkeys" and "Keypoint Detection for Injury Identification during Turkey Husbandry Using Neural Networks". The dataset contains cropped images of potential injuries which were separately annotated by three experts as not injured or injured. The Mice Bone dataset is based on the raw data which is available at https://doi.org/10.5281/zenodo.3355936 .The raw data are 3D scans from collagen fibers in mice bones. The three proposed classes are similar and dissimilar collagen fiber orientations and not relevant regions due to noise or background. We used the given segmentations to cut image regions from the original 2D image slices which mainly consist of one class. The CIFAR-10H dataset (https://github.com/jcpeterson/cifar-10h) provides multiple annotations for the test set of CIFAR-10. Technical description Each folder represents one dataset. The subfolders train, val and unlabeled represent the used data splits Training, Validation and Unlabeled respectively. The used ground-truth labels is given a folder name for each image. Each image is one datapoint. The filenames for the plankton data are the original Ecotaxa ID (https://ecotaxa.obs-vlfr.fr/).The filenames for the turkey data are a random number and the class 0 means not injured and class 1 is injured. The filesnames for the Mice Bone dataset are # # # .png. The filenames for the CIFAR-10H dataset are randomly generated counters. Each dataset has an additional dataset_import.json and an annotations.json file. The first file contain basically all information (image path, class names, datasplit and groundtruth class) as the file structured explained above. The second file contains the raw annotations per file. These annotations can be used to approximate the underlying ground truth distribution. The provided ground truth label was randomly selected from the approximation of the underyling ground truth based on this file. License Information CIFAR-10H is already published at https://github.com/jcpeterson/cifar-10h under Creative Commons BY-NC-SA 4.0 license. Full license information at https://creativecommons.org/licenses/by-nc-sa/4.0/ The CIFAR-10 image and original label data can be found at: https://www.cs.toronto.edu/~kriz/cifar.html The data was reformatted for this paper and is republished under Creative Commons BY-NC-SA 4.0 license. All other datasets (Plankton, Turkey, Mice Bone) are adapted works of previous publications. See above or below in the citation for the original works. The are republished under Creative Commons BY-SA 4.0 license. Full license information at https://creativecommons.org/licenses/by/4.0/ Citation Be aware that you need to reference this and previous works if you want to use this data. Please cite as: @Article{schmarje2022dc3, AUTHOR = {Schmarje, Lars and Santarossa, Monty and Schröder, Simon-Martin and Zelenka, Claudius and Kiko, Rainer and Stracke, Jenny and Volkmann, Nina and Koch, Reinhard}, TITLE = {A data-centric approach for improving ambiguous labels with combined semi-supervised classification and clustering}, JOURNAL = {Arxiv}, YEAR = {2022}, } Original data in @Article{Schmarje2021foc, AUTHOR = {Schmarje, Lars and Brünger, Johannes and Santarossa, Monty and Schröder, Simon-Martin and Kiko, Rainer and Koch, Reinhard}, TITLE = {Fuzzy Overclustering: Semi-Supervised Classification of Fuzzy Labels with Overclustering and Inverse Cross-Entropy}, JOURNAL = {Sensors}, VOLUME = {21}, YEAR = {2021}, NUMBER = {19}, ARTICLE-NUMBER = {6661}, URL = {https://www.mdpi.com/1424-8220/21/19/6661; https://doi.org/10.5281/zenodo.5550919}, ISSN = {1424-8220}, DOI = {10.3390/s21196661} } @article{peterson2019cifar10h, author = {Peterson, Joshua and Battleday, Ruairidh and Griffiths, Thomas and Russakovsky, Olga}, doi = {10.1109/ICCV.2019.00971}, eprint = {1908.07086}, isbn = {9781728148038}, issn = {15505499}, journal = {Proceedings of the IEEE International Conference on Computer Vision}, pages = {9616--9625}, title = {{Human uncertainty makes classification more robust}}, volume = {2019-Octob}, year = {2019} } @article{volkmann2021turkey, author = {Volkmann, Nina and Br{\"{u}}nger, Johannes and Stracke, Jenny and Zelenka, Claudius and Koch, Reinhard and Kemper, Nicole and Spindler, Birgit}, doi = {10.3390/ani11092655}, journal = {Animals 2021}, pages = {1--13}, title = {{Learn to train: Improving training data for a neural network to detect pecking injuries in turkeys}}, volume = {11}, year = {2021} } @article{schmarje2019, author = {Schmarje, Lars and Zelenka, Claudius and Geisen, Ulf and Gl{\"{u}}er, Claus-C. and Koch, Reinhard}, doi = {10.1007/978-3-030-33676-9_26}, eprint = {1907.12868}, isbn = {9783030336752}, issn = {23318422}, journal = {DAGM German Conference of Pattern Regocnition}, pages = {374--386}, publisher = {Springer}, title = {{2D and 3D Segmentation of uncertain local collagen fiber orientations in SHG microscopy}}, volume = {11824 LNCS}, year = {2019} }

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,004
score de la tête « metaresearch » (Gemma)0,014
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Jeu de données · Signal consensuel: Jeu de données
Score de désaccord entre enseignants0,022
Score d'incertitude au seuil0,074

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0040,014
Méta-épidémiologie (sens strict)0,0040,001
Méta-épidémiologie (sens large)0,0020,004
Bibliométrie0,0060,008
Études des sciences et des technologies0,0020,001
Communication savante0,0030,002
Science ouverte0,0050,004
Intégrité de la recherche0,0030,004
Charge utile insuffisante (le modèle a refusé de juger)0,0220,044

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,096
Tête enseignante GPT0,305
Écart entre enseignants0,209 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreJeu de données

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2022
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueZenodo (CERN European Organization for Nuclear Research)Même sujetAdvanced Clustering Algorithms ResearchTravaux en français237 207