Bibliographic record
Abstract
The Discrete Event System Specification (DEVS) provides a general methodology for hierarchical construction of reusable models in a modular way and has been used to simulate sophisticated systems in a variety of domains.This dissertation addresses software design and performance issues that arise in parallel simulation of large-scale DEVS-based models on both multiprocessor clusters and chip-multiprocessor architectures.The Time Warp (TW) mechanism is the most well-known optimistic synchronization protocol for Parallel Discrete-Event Simulations (PDES).With the increasing scale and complexity, TW simulations face new challenges in terms of excessive memory consumption and operational overhead.In an effort to alleviate these problems, a novel Lightweight Time Warp (LTW) protocol is proposed for efficient optimistic parallel DEVS simulation on multiprocessor clusters.By exploring the intrinsic computational properties of DEVS-based simulations, the LTW protocol allows purely optimistic parallel simulation to be driven by only a few full-fledged TW Logical Processes (LPs), while most of the LPs are set free from the burden of TW execution.The experimental results indicate that simulation performance can be improved significantly in various aspects, including shortened execution time, reduced memory footprint, lowered operational overhead, accelerated event queue operations, facilitated process migration, and enhanced system stability and scalability.To address the limitations of microprocessor performance, the industry is moving towards multicore chip-multiprocessor designs.As a latest example of this trend, the IBM Cell processor has attracted a growing interest from the modeling and simulation community.However, general-purpose PDES on such platform requires innovative redesign of existing algorithms in return for better simulation performance.To this end, a new computing technique called Multicore Acceleration of DEVS Systems (MADS) is developed for highperformance parallel DEVS simulation on the Cell processor, combining multi-grained parallelism and various optimizations to overcome the major performance bottlenecks, while hiding, to a great extent, the technical details of multicore programming from general users.Through the concept of LP virtualization, the MADS technique explicitly exploits the massive data-and event-level parallelism inherent in the simulation, making the achievable ?performance gain more deterministic and predictable than the traditional LP-oriented approaches.Promising results have been produced in the experiments, demonstrating that the MADS technique can be used to accelerate both memory-bound and compute-bound computational kernels in demanding parallel DEVS simulations.The proposed technique not only allows a broad community of DEVS users to tap the potential of the Cell processor with a minimal knowledge of the multicore execution environment, but also makes it possible to integrate cluster-based parallel simulation with multicore-accelerated parallel simulation on hybrid supercomputers.Vl two-level cache hierarchy (32KB Ll, 512KB L2) to access system main memory and provides top-level thread control for a parallel application, whereas each SPE can directly access only a private on-chip Local Storage (LS) of 256KB that contains both code and data (including the call stack) of an SPE thread.Data sharing is achieved mainly through software-managed, explicitly-addressed, autonomous Direct Memory Access (DMA) transfers, which require proper address alignment and transfer size to attain peak performance.The cores can also communicate 32-bit short messages via the on-chip Element Interconnect Bus (EIB) channels (e.g., mailboxes and signals).Moreover, the SPEs support both scalar and 128-bit SIMD operations that can be applied at 2, 4, 8, and 16way granularities.All these features make the Cell processor an attractive vehicle for studying new computing techniques on the emerging CMP architectures.On the flip side, the asymmetric design of heterogeneous cores with explicit memory control increases software complexity considerably and requires innovative redesign of existing algorithms to exploit parallelism at different system levels in return for better application performance.While the Cell processor is rapidly gaining popularity in scientific and multimedia applications (see, e.g., [Bad07a, Ged07, Pet07, and Sai07]), its potential has yet to be realized in PDES systems due to several challenging issues.First of all, most of the existing PDES techniques, developed with traditional parallel computing systems in mind, adopt a LP-oriented approach to partitioning a simulation across multiple nodes of a cluster [FujOO], while neglecting to integrate with other forms of parallelism (e.g., data-level parallelism, memory-level parallelism, and compute-transfer parallelism) that are made available on modern multicore platforms.As a result, developing efficient PDES algorithms on the Cell processor, and on CMP architectures in general, requires a holistic approach that takes into account all parallelization options provided by the processor microarchitecture.In addition, PDES programs typically involve highly irregular, control-intensive computation with complex data dependency and unpredictable memory access pattern [Fuj90], a class of workload that is generally regarded as not well-suited for parallelization on the Cell processor [Sca09a].Furthermore, recent advances towards facilitating software development on the Cell processor, in the form of compiler-assisted vectorization (e.g., [Eic06 and Kni07]) and middleware frameworks (e.g., [McC06 and Per07]), offer little help in parallelizing PDES systems, mainly because these techniques, applied at a lower software layer, lack the 6 1.2.2.Multicore Acceleration of DEVS Systems A novel computing technique, referred to as Multicore Acceleration of DEVS Systems (MADS), is proposed that combines multi-grained parallelism and various optimization strategies to overcome the major performance bottlenecks in demanding DEVS-based simulations on the Cell processor.The development of the MADS technique consists of the following contributions.• Two types of typical computational kernels are extracted from general-purpose DEVSbased simulations, reflecting the major performance bottlenecks in the simulation process, as illustrated by detailed simulation profiles.• The concept ofLP visualization is introduced to support flexible and efficient mapping of LPs to different processing elements of the Cell processor dynamically at runtime, improving the utilization of the heterogeneous cores, minimizing the synchronization overhead, and allowing for fine-grained dynamic load balancing.• Two forms of event-level parallelism are identified from a data-flow perspective, including the event-embarrassing parallelism and the event-streaming parallelism.Unlike the LP-oriented parallelization strategy adopted in most existing PDES systems, the MADS technique explicitly exploits the fine-grained event-level parallelism that is inherent in the DEVS simulation process, making the achievable parallelism more deterministic and predictable.• To accelerate the computational kernels, new simulation algorithms are developed to combine multi-grained parallelism at different levels of the system in a coherent way, including thread-level parallelism, data-level parallelism, event-level parallelism, datastreaming parallelism, and compute-I/O parallelism.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".