MétaCan
Menu
Retour à la cohorte
Enregistrement W4361806078 · doi:10.48550/arxiv.2304.00019

Workflows Community Summit 2022: A Roadmap Revolution

2023· preprint· en· W4361806078 sur OpenAlexaff
Rafael Ferreira da Silva, Rosa M. Badía, Venkat Bala, Debbie Bard, Peer‐Timo Bremer, Ian K. Buckley, Silvina Caíno‐Lores, Kyle Chard, Carole Goble, Shantenu Jha, Daniel S. Katz, Daniel Laney, Manish Parashar, Frédéric Suter, Thomas Uram, İlkay Altıntaş, Stefan Andersson, William Arndt, Juan Pedro Aznar, Jonathan Bader, Bartosz Baliś, Chris Blanton, Kelly Rosa Braghetto, Aharon Brodutch, Paul Brunk, Henri Casanova, Alba Cervera Lierta, Justin Chigu, Tainã Coleman, Nick Collier, Iacopo Colonnelli, Frederik Coppens, Michael R. Crusoe, W. S. Cunningham, Bruno de Paula Kinoshita, Paolo Di Tommaso, Charles Doutriaux, Matthew T. Downton, Wael Elwasif, Bjoern Enders, Christopher Erdmann, Thomas Fahringer, Ludmilla Figueiredo, Rosa Filgueira, Martin Foltín, Anne Fouilloux, Luiz Gadelha, Andy Gallo, Artur Garcia Saez, Daniel Garijo, Ryan E. Grant, Samuel Grayson, Patricia Grubel, Johan E. Gustafsson, Valérie Hayot‐Sasson, Óscar Hernández, Marcus Hilbrich, Annmary Justine, I. Laflotte, Fabian Lehmann, André Luckow, Jakob Luettgau, Motohiko Matsuda, Doriana Medić, Peter Mendygral, Marek T. Michalewicz, Jorji Nonaka, Maciej Pawlik, Loïc Pottier, Line Pouchard, Mathias Pütz, Santosh Kumar Radha, Lavanya Ramakrishnan, Sasko Ristov, Paul Romano, Daniel Rosendo, Martin Ruefenacht, Katarzyna Rycerz, Nishant Saurabh, V. Savchenko, Martin Schulz, Christine M. Simpson, Raül Sirvent, Tyler J. Skluzacek, Stian Soiland‐Reyes, Renan P. Souza, Sreenivas R. Sukumar, Ziheng Sun, Alan Sussman, Douglas Thain, Mikhail Titov, Benjamín Tovar, Aalap Tripathy, Matteo Turilli, Bartosz Tużnik, Hubertus J. J. van Dam, Aurelio Vivas, Logan Ward, Patrick Widener, Sean R. Wilkinson, Justyna Zawalska, Mahnoor Zulfiqar

Notice bibliographique

RevuearXiv (Cornell University) · 2023
Typepreprint
Langueen
DomaineDecision Sciences
ThématiqueScientific Computing and Data Management
Établissements canadiensQueen's UniversityAgnostiq (Canada)
Organismes subventionnairesOak Ridge National LaboratoryNational Nuclear Security AdministrationOffice of ScienceU.S. Department of Energy
Mots-clésWorkflowComputer scienceCloud computingData scienceCyberinfrastructureWorkflow management systemSoftware engineeringDatabaseOperating system

Résumé

récupéré en direct d'OpenAlex

Scientific workflows have become integral tools in broad scientific computing use cases. Science discovery is increasingly dependent on workflows to orchestrate large and complex scientific experiments that range from execution of a cloud-based data preprocessing pipeline to multi-facility instrument-to-edge-to-HPC computational workflows. Given the changing landscape of scientific computing and the evolving needs of emerging scientific applications, it is paramount that the development of novel scientific workflows and system functionalities seek to increase the efficiency, resilience, and pervasiveness of existing systems and applications. Specifically, the proliferation of machine learning/artificial intelligence (ML/AI) workflows, need for processing large scale datasets produced by instruments at the edge, intensification of near real-time data processing, support for long-term experiment campaigns, and emergence of quantum computing as an adjunct to HPC, have significantly changed the functional and operational requirements of workflow systems. Workflow systems now need to, for example, support data streams from the edge-to-cloud-to-HPC enable the management of many small-sized files, allow data reduction while ensuring high accuracy, orchestrate distributed services (workflows, instruments, data movement, provenance, publication, etc.) across computing and user facilities, among others. Further, to accelerate science, it is also necessary that these systems implement specifications/standards and APIs for seamless (horizontal and vertical) integration between systems and applications, as well as enabling the publication of workflows and their associated products according to the FAIR principles. This document reports on discussions and findings from the 2022 international edition of the Workflows Community Summit that took place on November 29 and 30, 2022.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,042
score de la tête « metaresearch » (Gemma)0,033
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Commentaire · Signal consensuel: Commentaire
Score de désaccord entre enseignants0,042
Score d'incertitude au seuil0,221

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0420,033
Méta-épidémiologie (sens strict)0,0020,001
Méta-épidémiologie (sens large)0,0010,002
Bibliométrie0,0030,003
Études des sciences et des technologies0,0040,004
Communication savante0,0090,015
Science ouverte0,0040,009
Intégrité de la recherche0,0090,010
Charge utile insuffisante (le modèle a refusé de juger)0,0250,019

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,434
Tête enseignante GPT0,295
Écart entre enseignants0,139 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreCommentaire

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations15
Publié2023
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revuearXiv (Cornell University)Même sujetScientific Computing and Data ManagementTravaux en français237 207