Analysis and monitoring of task allocation in DevOps processes
Bibliographic record
Abstract
In software development, the increasing need for rapid adaptation has led to the adoption of Agile methodologies, while the demand for continuous, high-quality deployment has driven the rise of DevOps practices. These approaches emphasize breaking work into smaller tasks, which improves team visibility and organization. However, task allocation—the process of assigning individuals to specific tasks—remains a persistent challenge. Effective allocation is crucial for maintaining efficiency and optimizing workflows. Although theoretical frameworks and algorithms exist, they often fail to align with industry realities, leaving organizations without reliable strategies. This thesis investigates task allocation in an industry-based context. Instead of focusing on advanced mathematical or machine learning models, it emphasizes deriving actionable insights closely aligned with real-world scenarios. The study introduces metrics and tools to analyze and improve task allocation. Conducted in collaboration with TELUS, a Canadian telecommunications provider, this study leverages real-world data from their GitLab projects. Regular communication with TELUS employees provided essential context, identified key areas for improvement, and validated hypotheses, ensuring relevance to both TELUS operations and broader industry practices. This research is based on the hypothesis that analyzing task lead time and identifying its influencing factors can improve task allocation. The methodology involves data extraction and processing, defining metrics to characterize tasks across multiple dimensions, and visualizing results. A key contribution is the definition of a criterion for characterizing tasks based on their ability to meet estimates, which serves as a foundation for identifying task failures and conducting related analyses. Analysis 3148 issues from 137 projects in TELUS revealed a consistent 34% task failure rate across estimated sizes and over time, underscoring the need for further tracking and improvement. The findings also indicate that the use of waiting columns raises the failure rate from 34% to 56%, a trend more pronounced in tasks with small estimates, highlighting the need for closer monitoring and analysis of their impact. Additionally, multitasking was analyzed as a potential contributor to task failure, but the findings were inconclusive, suggesting the need for further investigation. Moreover, 35% of failed tasks were underestimated, a trend particularly evident in smaller estimates—emphasizing the importance of improving task estimation processes. These insights provide project managers at TELUS and other companies with valuable tools to improve task planning and execution. This research presents a fully functional open-source framework for task allocation analysis, equipping organizations with a practical tool to identify inefficiencies. Rather than advocating for radical change, it prioritizes actionable insights and practical solutions for task assignment. These findings lay a strong foundation for future research, particularly in developing automated systems to detect and mitigate inefficient task allocation behaviors—offering valuable improvements for industry applications.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.067 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.007 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".