Understanding and circumventing deployed traffic differentiation practices
Bibliographic record
Abstract
Net neutrality, the principle that Internet Service Providers (ISPs) should treat all Internet communications equally, has been the subject of considerable public debate over the past decade. One specific example of a net neutrality violation is traffic differentiation: giving a better (or worse) performance to certain classes of Internet traffic. Despite the potential impact on content providers and users (e.g., a potential advantage to certain content providers but not others), there is little work that investigates current traffic differentiation practices. My work addresses this need by answering the following key questions: "How prevalent are traffic differentiation practices?","How are these policies implemented?", "What is the impact of these policies?", and "Is there an efficient way to circumvent them?" I argue that even without internal access to either content providers or ISPs, researchers can independently analyze traffic differentiation practices, and Internet users can evade middleboxes that ISPs commonly deploy for enforcing differentiation policies. Specifically, my work uncovers the current deployed traffic differentiation policies, analyzes how are differentiation policies implemented, and infers the impact of these practices on affected applications. With insight into the deployed practices, we evaluate opportunities to mitigate differentiation's impact and develop a system that can automatically circumvent middleboxes that enforce these policies. My work raises awareness of the prevalence of net neutrality violations and provides useful insights for both academia and the general public. More than 100,000 users contributed to the research; the work was covered by numerous media outlets, sparked collaborations with a regulator (Arcep), an ISP (Verizon), a content provider (Amazon AWS), and an open-source project (M-Lab); Legislators including Massachusetts state legislators, federal legislators, the FCC, the FTC, and CRTC in Canada cited my work when proposing new network regulations.--Author's abstract
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".