MétaCan
Menu
Back to cohort
Record W6983441807

Mise à l'échelle des analyses de trace dans une architecture modulaire

2024· other· fr· W6983441807 on OpenAlexfundno aff

Bibliographic record

VenuePolyPublie (École Polytechnique de Montréal) · 2024
Typeother
Languagefr
FieldAgricultural and Biological Sciences
TopicBerry genetics and cultivation research
Canadian institutionsnot available
FundersNatural Sciences and Engineering Research Council of Canada
KeywordsFilter (signal processing)NucleofectionLimitingTerm (time)
DOInot available

Abstract

fetched live from OpenAlex

RÉSUMÉ: Les systèmes de calcul de haute performance deviennent de plus en plus nécessaires pour répondre aux besoins auxquels nous sommes confrontés afin de fournir des services dont des millions d’utilisateurs dépendent chaque jour. Il est donc important de disposer d’outils axés sur la compréhension du fonctionnement du système, en collectant les informations nécessaires permettant de dresser des analyses. Dans ce contexte, le traçage, qui est une technique largement connue permettant de collecter les informations sur les fonctionnements internes et les états du système, constitue l’une des meilleures approches. Dans ce cas précis, l’information collectée nécessite le développement de visualisations basées sur des analyses définies en amont, et permettant l’interaction avec l’utilisateur dans son cycle de développement et d’opération. Il existe des solutions pour analyser les traces collectées dans les systèmes de calcul de haute performance, mais aucune d’entre elles ne résout complètement les besoins en évolutivité du traitement des analyses et en flexibilité du développement des analyses. Par conséquent, nous proposons une nouvelle solution qui permet la mise à l’échelle des analyses dans une architecture modulaire. Notre travail s’étend à l’implémentation d’une architecture distribuée visant à accroître l’évolutivité des outils d’analyse et visant à résoudre les problèmes de flexibilité lors du développement d’analyses. De plus, nous avons étendu le Trace Server Protocol. Cela nous a permis d’implémenter des analyses globales à l’échelle d’une grappe de calcul et faciliter l’analyse des traces connectées en exploitant les évènements réseaux. Des bancs de tests ont été menés afin d’évaluer le gain en performance et le surcoût associés à l’architecture distribuée Maître-Ouvrier par rapport à l’approche précédente, l’architecture distribuée client-serveur. En conclusion, les résultats montrent un gain en performance significatif en fonction du nombre de noeuds mis en place et indiquent un faible surcoût. La majorité du surcoût est causée par la sérialisation du format JSON. Nous avons conclu que la solution pourrait facilement être portée dans les systèmes de calcul de haute performance, mais aussi dans tout système distribué qui nécessite une puissance d’analyse de traces supérieure et une vaste gamme d’options d’analyse. Cependant, dans les cas où le volume des traces est faible, le gain pourrait être nul ou même se transformer en légère perte. ABSTRACT: High-performance computing systems are becoming increasingly necessary to meet the demands of services used by millions of users every day. Therefore, it is important to have tools focused on understanding system operation by collecting necessary information for analysis. In this context, tracing, a widely known technique for gathering information on internal operations and system states, emerges as one of the best approaches. In this specific case, the collected information requires the development of visualizations based on predefined analyses, enabling interaction with the user in their development. While there are solutions for analyzing traces collected in high-performance computing systems, none completely address the scalability needs of analysis processing and the flexibility of analysis development. Therefore, we propose a new solution that enables scalable analysis within a modular architecture. Our work extends to the implementation of a distributed architecture aimed at increasing the scalability of analysis tools. This modular distributed architecture is designed to address flexibility issues encountered during analysis development. It extends the Trace-Server-Protocol to enable global analysis at the scale of a computing cluster and to facilitate the analysis of connected traces by leveraging network events. We conducted a benchmark to evaluate the performance gain and associated overhead of the Master-Worker distributed architecture compared to the previous Client-Server distributed architecture. In conclusion, the results demonstrate a significant performance gain based on the number of nodes deployed and indicate a low overhead. The majority of the overhead is caused by JSON serialization. We concluded that the solution could easily be implemented in highperformance computing systems as well as any distributed system requiring enhanced trace analysis capabilities and a wide range of analysis options. However, in cases where trace volume is low, no performance or even a degradation may be experienced.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.013
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.014
Threshold uncertainty score0.027

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0040.013
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.001
Science and technology studies0.0010.001
Scholarly communication0.0040.004
Open science0.0010.002
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0050.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.030
GPT teacher head0.278
Teacher spread0.248 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2024
Admission routes1
Has abstractyes

Explore more

Same venuePolyPublie (École Polytechnique de Montréal)Same topicBerry genetics and cultivation researchFrench-language works237,207