DTraComp: Comparing distributed execution traces for understanding intermittent latency sources
Bibliographic record
Abstract
Microservice architectures can enhance software development by using multiple programming languages and deployment infrastructures, isolating failures within individual services, and accelerating the debugging and fixing of issues in independent services. Locating performance degradation becomes challenging, due to the presence of numerous service instances with complex interactions compounded by parallelism. Although end-to-end tracing allows tracing execution paths across services, and detecting their latencies, it is limited to high-level information. Indeed, end-to-end tracing cannot pinpoint the root causes of performance degradation between the processes. Moreover, many existing performance analysis tools lack a comparison feature to give developers a comprehensive view of the performance differences between two groups of requests. This paper introduces DTraComp (Distributed Trace Compare) , an open-source framework, compatible with various microservice trace standards, and integrated with Eclipse Trace Compass™. Our framework offers robust visual comparison capability for two groups of executions within distributed systems, which includes nested spans executed in parallel. Furthermore, it provides system kernel details for each thread involved in the execution of each span, allowing it to pinpoint the reasons for performance degradation across distributed systems. We used our proposed framework to analyze five practical use cases. By evaluating the efficiency of our tool, it was determined that the overall time complexity scales linearly O(n) with the trace size, indicating its suitability for deployment in production environments. It is currently used within Ericsson company for performance evaluation purposes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.010 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".