MétaCan
Menu
Back to cohort
Record W3115903506 · doi:10.1002/cpe.6153

Taming <scp>next‐generation HPC</scp> systems: <scp>Run‐time</scp> system and algorithmic advancements

2020· article· en· W3115903506 on OpenAlexaboutno aff
Roman Wyrzykowski, Bolesław K. Szymański

Bibliographic record

VenueConcurrency and Computation Practice and Experience · 2020
Typearticle
Languageen
FieldComputer Science
TopicDistributed and Parallel Computing Systems
Canadian institutionsnot available
Fundersnot available
KeywordsComputer scienceConcurrencyInformaticsOperations researchDistributed computingEngineering

Abstract

fetched live from OpenAlex

PPAM is a biennial series of international conferences dedicated to exchanging ideas between researchers involved in parallel and distributed computing, including theory and applications, as well as applied and computational mathematics.Twelve previous events have been held in different universities in Poland since 1994, when the first PPAM took place in Czestochowa.Thus, the event in Bialystok was an opportunity to celebrate the 25th anniversary of PPAM.The focus of PPAM 2019 was on models, algorithms, and software tools that facilitate efficient and convenient use of modern parallel and distributed computing systems, as well as on large-scale modern applications, including advances in machine learning and artificial intelligence.This meeting gathered more than 170 participants from 26 countries.The accepted papers were presented at the regular tracks of the PPAM 2019 conference and during the workshops.With each submission evaluated by at least three reviewers, a strict reviewing process resulted in the acceptance of 91 contributed papers for publication in the conference proceedings, while approximately 43% of the submissions were rejected.The Program Committee selected 41 papers for presentation in the regular conference track, resulting in an acceptance rate of about 46%.Based on the review results, 10 papers (11% of submissions) were selected for a special journal issue.Besides quality, another important criterion for selection was each paper's contribution to the thematic consistency of the issue.The focus of this special issue is on algorithmic advancements in matching the software properties to parallel architecture, including GPU accelerators and clusters.These advancements are crucial for successfully parallelizing such complex applications as simulating geophysical flows, solving ordinary differential equations (ODEs), structural analysis of nuclear reactor containment buildings, solving generalized eigenvalue problems, modeling of material science phenomena, and others.A complementary topic of this issue is advances in run-time systems since increasing levels of parallelism in multi-and many-core chips and the emerging heterogeneity of computational resources coupled with energy, resilience, and data movement constraints radically increase the importance of efficient run-time scheduling and execution control.After the conference, the Program Committee invited the authors of selected papers to submit revised and extended versions of their works.These new versions were reviewed independently again by at least three reviewers.Finally, nine contributions were accepted for publication.They are summarized below.Paper [1] focuses on the accurate assembly of the system matrix, which is an essential step in any code that solves partial differential equations on a mesh.This step can become costly in multigrid codes requiring cascades of matrices that depend upon each other, or dynamic adaptive mesh refinement.To reduce the time to solution, the authors propose that these constructions can be performed concurrently with the multigrid cycles.Furthermore, they desynchronize the assembly from the solution process.This non-trivial increase in the concurrency level improves the scalability.As assembly routines are notoriously memory-and bandwidth-demanding, the final algorithmic enhancement uses a hierarchical, lossy compression scheme that brings the memory footprint down aggressively even when the system matrix entries carry little information or are not yet available with high accuracy.An efficient algorithm for the parallel solution of indefinite saddle point systems with iterative solvers based on the Golub-Kahan bidiagonalization is presented in Reference [2].Such systems arise in many application fields, for example, in structural mechanics.A scalability study of the generalized solver shows improved performance for the two-dimensional (2D) Stokes equations compared to previous works.Furthermore, the authors investigate the performance of different parallel inner solvers in the outer Golub-Kahan iteration for a three-dimensional (3D) Stokes problem.When the number of cores is increasing for a fixed problem size, the solver exhibits good speedups of up to 50% with the 1024 cores.For the tests in which the problem size grows while the workload in each core stays constant, the performance of the solver scales almost linearly with the increase in the number of cores. Paper [3] proposes a locality optimization technique for the parallel solution on GPUs of large systems of ODEs by explicit one-step methods.This technique is based on tiling across the stages of a one-step method and is enabled by a special structure of the class of ODE systems-with the limited access distance.The paper focuses on increasing the range of access distances for which the tiling technique can provide a speedup

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.008
metaresearch head score (Gemma)0.022
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.031
Threshold uncertainty score0.103

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0080.022
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.004
Science and technology studies0.0020.004
Scholarly communication0.0090.015
Open science0.0030.006
Research integrity0.0020.005
Insufficient payload (model declined to judge)0.0310.008

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.042
GPT teacher head0.289
Teacher spread0.247 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2020
Admission routes1
Has abstractyes

Explore more

Same venueConcurrency and Computation Practice and ExperienceSame topicDistributed and Parallel Computing SystemsFrench-language works237,207