Health Data and Privacy in the Digital Era
Bibliographic record
Abstract
In 2010, the social networking site Facebook launched a platform allowing private companies to request users’ permission to access personal data. Few users were aware of the platform, which was integrated into Facebook’s terms of service. In 2014, Cambridge Analytica, a UK-based political consulting firm, developed a data-harvesting app. That app prompted Facebook users to provide psychological profiles, including responses such as “I get upset easily” and “I have frequent mood-swings” as part of a “research project.” The Facebook platform allowed users to share their friends’ data as well, enabling Cambridge Analytica to access tens of millions of personal profiles, identifying voters’ political preferences. The controversy revealed risks to identifiable health data posed by social media and web services companies’ practices. After the Cambridge Analytica controversy, Facebook suspended a project that aimed to link data about users’ medical conditions with information about their social networks. Individuals often reveal detailed, sensitive health information online. Through wearable devices, social media posts, traceable web searches, and online patient communities, users generate large volumes of health data. Although some individuals participate in online patient forums and wellness information sharing apps under their own names, others participate via pseudonyms, assuming their privacy is preserved. Many users believe their data will be shared only with those they designate.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.031 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.008 | 0.051 |
| Scholarly communication | 0.016 | 0.023 |
| Open science | 0.001 | 0.012 |
| Research integrity | 0.010 | 0.010 |
| Insufficient payload (model declined to judge) | 0.009 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".