Heavy Hitter Flow Detection using P4-based Programmable Data Plane Switches
Bibliographic record
Abstract
In data networks carrying large numbers of flows, Heavy Hitters (HHs) or Elephant flows are the flows exceeding pre-determined thresholds (e.g. wrt. number of packets or bytes) in a given time window. Such HH flows need to be handled differently in order to minimize their impact on other smaller flows. In recent years, HH detection techniques were shown to be effectively implementable in programmable data plane switches. In recent work, it was shown that the inter-packet gap can be used to identify heavy hitters. In order to conserve memory, such schemes use a limited-size hash table for storing flow state information and using this for the detection. However, when hash collisions occur, it is possible that a valid HH flow in the table can be replaced by a non-HH flow resulting in missing detection of HH flows. To address this problem, this paper incorporates a flow’s medium term Packet Count feature. In order to limit the packet count field size in the hash table, counting is done only till Hash Collision occurs so as to the reduce the range of values to be stored and thus, the required number of bits. The proposed scheme has been implemented in the P4 language and run on Intel Tofino hardware. Performance evaluation has been done using MAWI-based real-life traffic traces. The results shows that in several scenarios cases, we can significantly reduce the False Negatives for HHs by using the packet count data effectively and efficiently.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".