MétaCan
Menu
Retour à la cohorte
Enregistrement W3115903506 · doi:10.1002/cpe.6153

Taming <scp>next‐generation HPC</scp> systems: <scp>Run‐time</scp> system and algorithmic advancements

2020· article· en· W3115903506 sur OpenAlexaboutno aff
Roman Wyrzykowski, Bolesław K. Szymański

Notice bibliographique

RevueConcurrency and Computation Practice and Experience · 2020
Typearticle
Langueen
DomaineComputer Science
ThématiqueDistributed and Parallel Computing Systems
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésComputer scienceConcurrencyInformaticsOperations researchDistributed computingEngineering

Résumé

récupéré en direct d'OpenAlex

PPAM is a biennial series of international conferences dedicated to exchanging ideas between researchers involved in parallel and distributed computing, including theory and applications, as well as applied and computational mathematics.Twelve previous events have been held in different universities in Poland since 1994, when the first PPAM took place in Czestochowa.Thus, the event in Bialystok was an opportunity to celebrate the 25th anniversary of PPAM.The focus of PPAM 2019 was on models, algorithms, and software tools that facilitate efficient and convenient use of modern parallel and distributed computing systems, as well as on large-scale modern applications, including advances in machine learning and artificial intelligence.This meeting gathered more than 170 participants from 26 countries.The accepted papers were presented at the regular tracks of the PPAM 2019 conference and during the workshops.With each submission evaluated by at least three reviewers, a strict reviewing process resulted in the acceptance of 91 contributed papers for publication in the conference proceedings, while approximately 43% of the submissions were rejected.The Program Committee selected 41 papers for presentation in the regular conference track, resulting in an acceptance rate of about 46%.Based on the review results, 10 papers (11% of submissions) were selected for a special journal issue.Besides quality, another important criterion for selection was each paper's contribution to the thematic consistency of the issue.The focus of this special issue is on algorithmic advancements in matching the software properties to parallel architecture, including GPU accelerators and clusters.These advancements are crucial for successfully parallelizing such complex applications as simulating geophysical flows, solving ordinary differential equations (ODEs), structural analysis of nuclear reactor containment buildings, solving generalized eigenvalue problems, modeling of material science phenomena, and others.A complementary topic of this issue is advances in run-time systems since increasing levels of parallelism in multi-and many-core chips and the emerging heterogeneity of computational resources coupled with energy, resilience, and data movement constraints radically increase the importance of efficient run-time scheduling and execution control.After the conference, the Program Committee invited the authors of selected papers to submit revised and extended versions of their works.These new versions were reviewed independently again by at least three reviewers.Finally, nine contributions were accepted for publication.They are summarized below.Paper [1] focuses on the accurate assembly of the system matrix, which is an essential step in any code that solves partial differential equations on a mesh.This step can become costly in multigrid codes requiring cascades of matrices that depend upon each other, or dynamic adaptive mesh refinement.To reduce the time to solution, the authors propose that these constructions can be performed concurrently with the multigrid cycles.Furthermore, they desynchronize the assembly from the solution process.This non-trivial increase in the concurrency level improves the scalability.As assembly routines are notoriously memory-and bandwidth-demanding, the final algorithmic enhancement uses a hierarchical, lossy compression scheme that brings the memory footprint down aggressively even when the system matrix entries carry little information or are not yet available with high accuracy.An efficient algorithm for the parallel solution of indefinite saddle point systems with iterative solvers based on the Golub-Kahan bidiagonalization is presented in Reference [2].Such systems arise in many application fields, for example, in structural mechanics.A scalability study of the generalized solver shows improved performance for the two-dimensional (2D) Stokes equations compared to previous works.Furthermore, the authors investigate the performance of different parallel inner solvers in the outer Golub-Kahan iteration for a three-dimensional (3D) Stokes problem.When the number of cores is increasing for a fixed problem size, the solver exhibits good speedups of up to 50% with the 1024 cores.For the tests in which the problem size grows while the workload in each core stays constant, the performance of the solver scales almost linearly with the increase in the number of cores. Paper [3] proposes a locality optimization technique for the parallel solution on GPUs of large systems of ODEs by explicit one-step methods.This technique is based on tiling across the stages of a one-step method and is enabled by a special structure of the class of ODE systems-with the limited access distance.The paper focuses on increasing the range of access distances for which the tiling technique can provide a speedup

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,008
score de la tête « metaresearch » (Gemma)0,022
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Simulation ou modélisation · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: aucune
Score de désaccord entre enseignants0,031
Score d'incertitude au seuil0,103

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0080,022
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0020,004
Études des sciences et des technologies0,0020,004
Communication savante0,0090,015
Science ouverte0,0030,006
Intégrité de la recherche0,0020,005
Charge utile insuffisante (le modèle a refusé de juger)0,0310,008

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,042
Tête enseignante GPT0,289
Écart entre enseignants0,247 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSimulation ou modélisation
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations1
Publié2020
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueConcurrency and Computation Practice and ExperienceMême sujetDistributed and Parallel Computing SystemsTravaux en français237 207