The internet on its way back to a future of human dignity?: One year EU General Data Protection Regulation
Bibliographic record
Abstract
The General Data Protection Regulation (GDPR) has come a long way since it was first tabled as proposal by the European Commission on 25 January 2012. Probably, it will be remembered as the biggest achievement of the outgoing Juncker Commission. Despite the fact that Europe plays a second-tier role when it comes to the development and production of data-driven technology and services, research shows that of 132 countries which have national data protection laws at the beginning of 2019 the majority adopts the European ‘omnibus approach’. In other words, the protection of personal data has been promoted from a niche issue that is addressed on case by case basis (the ‘sectoral approach’ which is typical for US regulation) to a universal concern. Even more, privacy in the digital age is a popular human rights issue deserving the care of a ‘dedicated ambassador’ (Special Rapporteur) of the United Nations. Since becoming enforceable, GDPR has been proposed as sort of a gold standard to the international community and people across the world are happy to be ‘protected’ by it – even outside Europe. Corporations such as Facebook have gone from stating that ‘privacy is dead’ in 2010 to boldly claim that ‘the future is private’ in 2019. However, despite its undeniable popularity and societal impact it seems wrong to hail GDPR as the finest of regulatory instruments. As I will aim to demonstrate in this piece there remain important issues for which GDPR is an unsatisfying solution. As the digital layer of societal interaction evolves it will be crucial to take additional measures in the near term. At least, if the goal is to avoid fragmentation of the digital sphere, which for Europeans would result in limitation to the bubble of a ‘bourgeouis internet’.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.024 | 0.036 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.005 | 0.009 |
| Scholarly communication | 0.027 | 0.019 |
| Open science | 0.004 | 0.010 |
| Research integrity | 0.063 | 0.028 |
| Insufficient payload (model declined to judge) | 0.007 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".