MétaCan
Menu
Retour à la cohorte
Enregistrement W4233379804 · doi:10.22215/etd/2010-09753

Algorithms for parallel simulation of large-scale DEVS and cell-DEVS models

2010· dissertation· en· W4233379804 sur OpenAlexaff
Qi Liu

Notice bibliographique

Revuenon disponible
Typedissertation
Langueen
DomaineDecision Sciences
ThématiqueSimulation Techniques and Applications
Établissements canadiensCarleton UniversityCanadian HeritageLibrary and Archives Canada
Organismes subventionnairesnon disponible
Mots-clésDEVSComputer scienceScale (ratio)Parallel computingDiscrete event simulationComputational scienceModeling and simulationSimulationGeographyCartography

Résumé

récupéré en direct d'OpenAlex

The Discrete Event System Specification (DEVS) provides a general methodology for hierarchical construction of reusable models in a modular way and has been used to simulate sophisticated systems in a variety of domains.This dissertation addresses software design and performance issues that arise in parallel simulation of large-scale DEVS-based models on both multiprocessor clusters and chip-multiprocessor architectures.The Time Warp (TW) mechanism is the most well-known optimistic synchronization protocol for Parallel Discrete-Event Simulations (PDES).With the increasing scale and complexity, TW simulations face new challenges in terms of excessive memory consumption and operational overhead.In an effort to alleviate these problems, a novel Lightweight Time Warp (LTW) protocol is proposed for efficient optimistic parallel DEVS simulation on multiprocessor clusters.By exploring the intrinsic computational properties of DEVS-based simulations, the LTW protocol allows purely optimistic parallel simulation to be driven by only a few full-fledged TW Logical Processes (LPs), while most of the LPs are set free from the burden of TW execution.The experimental results indicate that simulation performance can be improved significantly in various aspects, including shortened execution time, reduced memory footprint, lowered operational overhead, accelerated event queue operations, facilitated process migration, and enhanced system stability and scalability.To address the limitations of microprocessor performance, the industry is moving towards multicore chip-multiprocessor designs.As a latest example of this trend, the IBM Cell processor has attracted a growing interest from the modeling and simulation community.However, general-purpose PDES on such platform requires innovative redesign of existing algorithms in return for better simulation performance.To this end, a new computing technique called Multicore Acceleration of DEVS Systems (MADS) is developed for highperformance parallel DEVS simulation on the Cell processor, combining multi-grained parallelism and various optimizations to overcome the major performance bottlenecks, while hiding, to a great extent, the technical details of multicore programming from general users.Through the concept of LP virtualization, the MADS technique explicitly exploits the massive data-and event-level parallelism inherent in the simulation, making the achievable ?performance gain more deterministic and predictable than the traditional LP-oriented approaches.Promising results have been produced in the experiments, demonstrating that the MADS technique can be used to accelerate both memory-bound and compute-bound computational kernels in demanding parallel DEVS simulations.The proposed technique not only allows a broad community of DEVS users to tap the potential of the Cell processor with a minimal knowledge of the multicore execution environment, but also makes it possible to integrate cluster-based parallel simulation with multicore-accelerated parallel simulation on hybrid supercomputers.Vl two-level cache hierarchy (32KB Ll, 512KB L2) to access system main memory and provides top-level thread control for a parallel application, whereas each SPE can directly access only a private on-chip Local Storage (LS) of 256KB that contains both code and data (including the call stack) of an SPE thread.Data sharing is achieved mainly through software-managed, explicitly-addressed, autonomous Direct Memory Access (DMA) transfers, which require proper address alignment and transfer size to attain peak performance.The cores can also communicate 32-bit short messages via the on-chip Element Interconnect Bus (EIB) channels (e.g., mailboxes and signals).Moreover, the SPEs support both scalar and 128-bit SIMD operations that can be applied at 2, 4, 8, and 16way granularities.All these features make the Cell processor an attractive vehicle for studying new computing techniques on the emerging CMP architectures.On the flip side, the asymmetric design of heterogeneous cores with explicit memory control increases software complexity considerably and requires innovative redesign of existing algorithms to exploit parallelism at different system levels in return for better application performance.While the Cell processor is rapidly gaining popularity in scientific and multimedia applications (see, e.g., [Bad07a, Ged07, Pet07, and Sai07]), its potential has yet to be realized in PDES systems due to several challenging issues.First of all, most of the existing PDES techniques, developed with traditional parallel computing systems in mind, adopt a LP-oriented approach to partitioning a simulation across multiple nodes of a cluster [FujOO], while neglecting to integrate with other forms of parallelism (e.g., data-level parallelism, memory-level parallelism, and compute-transfer parallelism) that are made available on modern multicore platforms.As a result, developing efficient PDES algorithms on the Cell processor, and on CMP architectures in general, requires a holistic approach that takes into account all parallelization options provided by the processor microarchitecture.In addition, PDES programs typically involve highly irregular, control-intensive computation with complex data dependency and unpredictable memory access pattern [Fuj90], a class of workload that is generally regarded as not well-suited for parallelization on the Cell processor [Sca09a].Furthermore, recent advances towards facilitating software development on the Cell processor, in the form of compiler-assisted vectorization (e.g., [Eic06 and Kni07]) and middleware frameworks (e.g., [McC06 and Per07]), offer little help in parallelizing PDES systems, mainly because these techniques, applied at a lower software layer, lack the 6 1.2.2.Multicore Acceleration of DEVS Systems A novel computing technique, referred to as Multicore Acceleration of DEVS Systems (MADS), is proposed that combines multi-grained parallelism and various optimization strategies to overcome the major performance bottlenecks in demanding DEVS-based simulations on the Cell processor.The development of the MADS technique consists of the following contributions.• Two types of typical computational kernels are extracted from general-purpose DEVSbased simulations, reflecting the major performance bottlenecks in the simulation process, as illustrated by detailed simulation profiles.• The concept ofLP visualization is introduced to support flexible and efficient mapping of LPs to different processing elements of the Cell processor dynamically at runtime, improving the utilization of the heterogeneous cores, minimizing the synchronization overhead, and allowing for fine-grained dynamic load balancing.• Two forms of event-level parallelism are identified from a data-flow perspective, including the event-embarrassing parallelism and the event-streaming parallelism.Unlike the LP-oriented parallelization strategy adopted in most existing PDES systems, the MADS technique explicitly exploits the fine-grained event-level parallelism that is inherent in the DEVS simulation process, making the achievable parallelism more deterministic and predictable.• To accelerate the computational kernels, new simulation algorithms are developed to combine multi-grained parallelism at different levels of the system in a coherent way, including thread-level parallelism, data-level parallelism, event-level parallelism, datastreaming parallelism, and compute-I/O parallelism.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,001
score de la tête « metaresearch » (Gemma)0,005
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Simulation ou modélisation · Signal consensuel: Simulation ou modélisation
GenreSignal candidat: Empirique · Signal consensuel: aucune
Score de désaccord entre enseignants0,017
Score d'incertitude au seuil0,035

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0010,005
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0010,001
Études des sciences et des technologies0,0010,001
Communication savante0,0010,001
Science ouverte0,0020,002
Intégrité de la recherche0,0020,002
Charge utile insuffisante (le modèle a refusé de juger)0,0060,001

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,098
Tête enseignante GPT0,439
Écart entre enseignants0,341 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSimulation ou modélisation
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations4
Publié2010
Routes d'admission1
Résumé présentoui

Explorer davantage

Même sujetSimulation Techniques and ApplicationsTravaux en français237 207