Architecture of a social media bot detection system
Bibliographic record
Abstract
Modern information systems require efficient architectures to ensure high performance, scalability, and reliability. This article presents an approach to system architecture design that incorporates the latest technological solutions and methods for optimizing the processing of large data sets. The paper proposes an original architecture of a bot detection system based on the microservices paradigm and modern data processing techniques. Unlike existing solutions, the proposed system does not aim to develop a radically new classification method but focuses on the effective integration of well-established approaches within a unified architecture. The advancement of information technologies requires the development of architectural solutions that guarantee high performance and reliability of software systems. With the increasing volume of data and growing demands for processing speed, traditional architectural approaches require refinement. Research in this field is important for software developers and system architects. The aim of this study is to develop an architectural concept that meets modern requirements for performance, scalability, and security. The main objectives include analyzing existing approaches, identifying their advantages and drawbacks, and designing an efficient architecture that minimizes resource consumption and increases data processing speed. The study employed methods of architectural analysis, system modeling, performance testing, and comparative evaluation of different approaches. For the implementation of the architecture, modern technologies were used, including the microservices paradigm, containerization, and distributed computing. The proposed architecture improves system performance by optimizing request processing and distributing workloads across services. The use of containerization and orchestration enables flexible scalability and enhances system stability. Performance analysis has shown reduced request processing latency and efficient utilization of server resources. The developed architecture has proven its effectiveness in test environments and can be applied to high-load systems. Future research directions include the integration of artificial intelligence for automatic scaling and service optimization, as well as studying the impact of different caching strategies on overall system performance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".