MétaCan
Menu
Retour à la cohorte
Enregistrement W6961096523 · doi:10.14466/cefasdatahub.126

A data product derived from Northeast Atlantic groundfish data from scientific trawl surveys 1983-2020

2022· dataset· en· W6961096523 sur OpenAlexaboutno aff

Notice bibliographique

RevueCefas · 2022
Typedataset
Langueen
DomaineBiochemistry, Genetics and Molecular Biology
ThématiqueGene expression and cancer classification
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésGroundfishFishingQuarter (Canadian coin)Sampling (signal processing)BayBycatchSurvey data collectionEnvironmental dataSurvey methodology

Résumé

récupéré en direct d'OpenAlex

This is a data product to support state indicators that are based from groundfish biological data, derived using primary data from surveys undertaken in the Northeast Atlantic between 1983 and 2020. Catch records by taxonomic group and by length category in terms of biomass and numbers of fish standardised to duration (per hour) or to the area swept by the haul. Data are available from multiple surveys using data downloaded from the ICES database of trawl surveys (DATRAS) once quality-controlled and standardised following procedures detailed in Greenstreet and Moriarty 2017. Data file names reflect the OSPAR region sampled, country conducting the sampling, fishing gear and time of years of sampling (as defined by Greenstreet and Moriarty 2017), e.g.: BBICFraBT4 refers to Bay of Biscay and Iberian Coast data from France by a Beam Trawl survey in quarter 4 of the year and GNSIntOT3 refers to Greater North Sea data from International (multiple countries) sampling by an Otter Trawl survey in quarter 3 of the year etc. Greenstreet, S.P.R. and Moriarty, M. (2017) OSPAR Interim Assessment 2107 Fish Indicator Data Manual (Relating to Version 2 of the Groundfish Survey Monitoring and Assessment Data Product). Scottish Marine and Freshwater Science Vol 8 No 17, 83pp. DOI: 10.7489/1985-1 Scientific survey data collected by multiple countries and made available through ICES DATRAS (https://www.ices.dk/data/data-portals/Pages/DATRAS.aspx). Swept-area estimates were generated by ICES 2021ab (ICES. 2021a. Workshop on the production of swept-area estimates for all hauls in DATRAS for biodiversity assessments (WKSAE-DATRAS). ICES Scientific Reports. 3:74. https://doi.org/10.17895/ices.pub.8232; ICES. 2021b. Workshop on the production of abundance estimates for sensitive species (WKABSENS); ICES Scientific Reports. 3:96. https://doi.org/10.17895/ices.pub.8299). ICES Data Centre host the database of trawl surveys (DATRAS) for groundfish and beam trawl data. DATRAS has an integrated quality check utility. All data, before entering the database, have to pass an extensive quality check. Despite this errors and missing data arise, which are subsequently dealt with by the data submitters from the contributing countries as required. However, this screening process was implemented in 2009 for data from 2004 onwards. Since some survey time-series extend back to the 1960s, historic data (unless re-evaluated and re-submitted by contributing countries) may not have been subject to the same level of quality control as these more recent data. Furthermore, the type of information collected, the level of detail and resolution in the data, has gradually evolved over time. In order to derive a single format, quality assured monitoring programme data product covering the entire Northeast Atlantic region inconsistencies in the datasets required resolution. These corrections are detailed in ICES 2021a,b: Biological data for trawl surveys are downloaded directly from DATRAS in raw exchange format (known as “HL data”). Ancillary data were processed by ICES 2021a,b to create the “SweptAreaAssessmentOutput” (which replaces the “HH data”) and these were downloaded from the same location: https://datras.ices.dk/Data_products/Download/Download_Data_public.aspx Data are processed to create a standalone data product to be used for indicator assessments of fish and elasmobranchs. Initially, hauls are subset to determine the Standard Monitoring Programme (i.e. excluding invalid hauls including those of duration shorter than 13 minutes or longer than 66 minutes, following Greenstreet and Moriarty 2017) and these hauls are used to define the Standard Survey Area by excluding areas sampled infrequently over time). Biological data were accepted with ICES SpecVal of 1, 4, 7, 10 (see http://vocab.ices.dk/ for further information on SpecVal categories). Additional QA/QC is made at this step to determine if species identification issues are present in the raw biological data and these were discussed and agreed with the Chief Scientist for each survey.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,002
score de la tête « metaresearch » (Gemma)0,008
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: aucune
GenreSignal candidat: Jeu de données · Signal consensuel: Jeu de données
Score de désaccord entre enseignants0,231
Score d'incertitude au seuil0,460

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0020,008
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0060,012
Études des sciences et des technologies0,0000,000
Communication savante0,0020,002
Science ouverte0,0010,001
Intégrité de la recherche0,0000,001
Charge utile insuffisante (le modèle a refusé de juger)0,0440,027

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,071
Tête enseignante GPT0,300
Écart entre enseignants0,229 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeObservationnel
Domainenon disponible
GenreJeu de données

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations3
Publié2022
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueCefasMême sujetGene expression and cancer classificationTravaux en français237 207