Black Hole Prediction in Backbone Networks: A Comprehensive and Type-Independent Forecasting Model
Bibliographic record
Abstract
Network backbone black holes(BH) pose significant challenges in the Internet by causing disruptions and data loss as routers silently drop packets without notification. These silent BH failures, stemming from issues like hardware malfunctions or misconfigurations, uniquely affect point-to-point packet flows without disrupting the entire network. Unlike cyber attacks and network intrusions, BHs are often untraceable, making early detection vital and challenging. This study addresses the need for an effective forecasting solution for BH occurrences, especially in environments with unlabeled traffic data where traditional anomaly detection methods fall short. The Type-Independent Black Hole Forecasting Model is introduced to predict BH occurrences with high precision across various anomalies, including contextual and collective anomaly types. The three-stage methodology processes unlabeled time-series network data, where the data is not pre-labeled as anomaly or normal, using machine learning and deep learning techniques to identify and forecast potential BH occurrences. The ’Point BH Identification and Segregation’ stage segregates point BH traffic using Density-Based Spatial Clustering of Applications with Noise(DBSCAN), followed by Reintegration and Time Series Smoothing. The final stage, Advanced Contextual and Collective BH Detection leverages Convolutional AutoEncoder(Conv-AE) with window sliding for advanced anomaly detection. Evaluation using a dual-dataset approach, including real backbone network traffic and a time-series adapted public dataset, demonstrates the adaptability of the model to real backbone BH detection systems. Experimental results show superior performance compared to state-of-the-art unsupervised anomaly forecasting models, with a 98% detection rate and 90% F-1 score, outperforming models like MultiHeadSelfAttention, which is the main building block of Transformers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".