Big Data and surveillance: Hype, commercial logics and new intimate spheres
Bibliographic record
Abstract
Big Data Analytics promises to help companies and public sector service providers anticipate consumer and service user behaviours so that they can be targeted in greater depth. The attempts made by these organisations to connect analytically with users raise questions about whether surveillance, and its associated ethical and rights-based concerns, are intensified. The articles in this special themed issue explore this question from both organisational and user perspectives. They highlight the hype which firms use to drive consumer, employee and service user engagement with analytics within both private and public spaces. Further, they explore extent to which, through Big Data, there is an attempt to expand surveillance into the emotional registers of domestic, embodied experience. Collectively, the papers reveal a fascinating nexus between the much-vaunted potential of analytics, the data practices themselves and the newly configured intimate spheres which have been drawn into the commercial value chain. Together, they highlight the need for conceptual and regulatory innovation so that analytics in practice may be better understood and critiqued. Whilst there is now a rich variety of scholarship on Big Data Analytics, critical perspectives on the organising practices of Big Data Analytics and its surveillance implications are thin on the ground. Combined, the articles published in this special theme begin to address this shortcoming.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.027 | 0.037 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.006 |
| Science and technology studies | 0.011 | 0.113 |
| Scholarly communication | 0.042 | 0.058 |
| Open science | 0.002 | 0.018 |
| Research integrity | 0.007 | 0.018 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".