Identificación de ataques de denegación de servicio distribuido (DDoS) mediante la integración de algoritmos de aprendizaje automático y arquitecturas de redes neuronales artificiales.
Bibliographic record
Abstract
Objective: To identify distributed denial of service (DDoS) attacks by integrating machine learning algorithms and artificial neural network architectures. Methodology: To structure the data analysis, the Knowledge Discovery Data (KDD) technique is used. This approach allows examining large volumes of information of various types, with the objective of identifying patterns, correlations and producing valuable information. As for the data set, the CIC-DDoS2019 dataset developed by the Canadian Cybersecurity Institute is used. Results: When training and evaluating the different algorithms, it was observed that the models based on decision trees, such as Random Forest and XGBoost, stood out for achieving the best results in terms of accuracy and efficiency. On the other hand, in the analysis of the performance of the neural networks, the Closed Stream Units (GRU) stood out by obtaining the best results in accuracy and precision. This performance suggests that GRUs achieve an optimal balance between predictive ability and minimization of false positives and negatives. Discussion: In the comparison between traditional machine learning models and neural networks for DDoS attack detection, it is observed that algorithms such as XGBoost and Random Forest offer similar or superior performance in terms of accuracy and also exhibit significantly shorter execution times. On the other hand, neural networks such as GRU and RNN achieve high accuracy, but with a high computational cost. Conclusions: XGBoost, demonstrated an optimal balance between accuracy (F1-score: 0.9992) and speed (11.47s), positioning itself as the most viable alternative for real-time implementations. In the field of neural networks, Gated Stream Units (GCU) obtained the best performance (accuracy: 0.9992; F1-score: 0.9992), given the ability to process temporal dependencies and reduce false positives.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".