The Development of the Data System and Growth in Data Sharing
Notice bibliographique
Résumé
A great wealth of ocean data exists, for a wide range of disciplines, derived from in-situ and remote sensing observing platforms, in real-time, near-real-time and delayed mode. These data are acquired as part of routine monitoring activities and as part of scientific surveys by a few thousand institutes and agencies all around the world. Both the means to acquire these data and the way in which they are used have changed greatly in the past ten years. Over the last decade, information technology has progressed a great deal. It presently allows the exchange of gigabytes of data and more via the Internet in developed countries. In the late nineties, it was considered high technology to provide data on CDROM rather than on magnetic tapes and only small datasets were distributed via the Internet. Nowadays CDROMs are considered as a backup delivery system especially for countries with poor Internet connections. The explosion in use of the Internet has provided new communications capabilities, new tools, and a new way of using computers. The nature of requirements from government agencies have changed: they want to know or estimate what the future of the earth will look like and what will be the impact on their territories due to climate change issues, ocean health monitoring and fisheries assessment, but they can't pay the full bill for the data acquisition. Therefore, they are pushing, and nowadays more often imposing, a change in data policy and a move towards increased data sharing, in which data acquired with public funds should be freely available to the community. Moreover, the nature of science itself has changed. Investigators and research funding agencies are looking for context, impacts, and synthesis, rather than just focusing on individual, well-defined processes. Most scientists need the data collected by others as well as their own. They cannot do their work using only data they have collected themselves. Finally, the growth of operational oceanographic services, based on downscaling of global model results, is really important. These are in demand by users especially for real/near real-time data just as operational meteorology has been doing for a long time.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,038 | 0,072 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,002 |
| Bibliométrie | 0,007 | 0,015 |
| Études des sciences et des technologies | 0,003 | 0,007 |
| Communication savante | 0,014 | 0,040 |
| Science ouverte | 0,008 | 0,018 |
| Intégrité de la recherche | 0,005 | 0,010 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,018 | 0,012 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».