Implementación de un servicio de detección de ataques de denegación de sistemas
Bibliographic record
Abstract
Los ataques DDoS (Denial of Service) han sido una amenaza persistente en el ciberespacio desde la década de 1990, y su sofisticación y frecuencia han ido en aumento. Estos ataques implican la saturación maliciosa de recursos de red o sistemas informáticos, lo que resulta en la interrupción del servicio para usuarios legítimos. Esta actividad no solo representa un problema económico para las empresas, sino que, cuando los objetivos son infraestructuras críticas como las de comunicaciones o sanitarias, estos ataques pueden utilizarse como métodos de desestabilización. \n \nEn la actualidad, discernir si una conexión es legítima o no es complicado. Los sistemas pueden enfrentar dificultades para manejar el volumen del tráfico generado en estos ataques y tienden a generar falsos positivos, bloqueando así conexiones legítimas. El objetivo de este trabajo es recopilar datos, limpiarlos y prepararlos, desarrollar un modelo de deep learning capaz de detectar estas conexiones y luego implementar y desplegar el modelo en la nube pública. \n \nEn la primera parte del trabajo, se exploran los conjuntos de datos CIC-DDoS2019 y CIC-IDS2017 del Canadian Institute for Cybersecurity. Se eliminan y corrigen valores atípicos y se seleccionan variables utilizando técnicas como Random Forest y la importancia de variables para reducir la dimensionalidad del conjunto de datos. Posteriormente, se implementa y entrena una red FNN siguiendo la arquitectura ResNet. \n \nEn la segunda sección, se desarrolla un panel de control que permite visualizar un informe de las conexiones después de pasar por el modelo. Este panel se despliega en la nube de AWS, junto con un backend que proporciona información a la página web y otros mecanismos para la limpieza de datos y la recopilación en tiempo real. \n \nAbstract: \n \nDDoS (Denial of Service) attacks have been a persistent threat in cyberspace since the 1990s, with their sophistication and frequency steadily increasing. These attacks involve the malicious saturation of network resources or computer systems, resulting in service disruption for legitimate users. This activity not only poses an economic problem for businesses but when the targets are critical infrastructures such as communication or healthcare systems, these attacks can be used as destabilization methods. \n \nCurrently, discerning whether a connection is legitimate or not is challenging. Systems may struggle to handle the volume of traffic generated in these attacks and tend to produce false positives, thus blocking legitimate connections. The objective of this work is to collect data, clean and prepare it, develop a deep learning model capable of detecting these connections, and then implement and deploy the model in the public cloud. \n \nIn the first part of the work, datasets such as CIC-DDoS2019 and CIC-IDS2017 from the Canadian Institute for Cybersecurity are explored. Outliers are removed and corrected, and variables are selected using techniques like Random Forest and variable importance to reduce the dimensionality of the dataset. Subsequently, a FNN network following the ResNet architecture is implemented and trained. \n \nIn the second section, a control panel is developed that allows visualization of a report of connections after passing through the model. This panel is deployed on the AWS cloud, along with a backend that provides information to the webpage and other mechanisms for real-time data cleaning and collection.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".