Personalization, analytics, and sponsored services: The challenges of applying PIPEDA to online tracking and profiling activities
Bibliographic record
Abstract
In 2008, the online advertising industry was found to be worth 27 billion dollars, a figure that was projected to double over the subsequent four years.1 The reason for this extraordinary market growth can be explained by two factors. To begin with, current technology now makes it possible to gather a great variety of information associated with a particular device or individual, including browsing history, which can be used to create a profile specific to that device or individual. This practice facilitates more personalized advertising, tailored to the interests and tastes of the consumer. Secondly, many online services, in the form of information or entertainment, are offered for free to consumers as long as they accept the presence of advertising and the eventuality that their online behaviour will be tracked to a certain degree. Internet business models are increasingly being based on the notion of greater customization of services and products. This entails that there are huge amounts of data that need to be collected about online users. Moreover, online profiles present new types of concerns. For instance, although isolated pieces of profile information may not be sensitive, their context, especially in light of profiling or behavioural analysis practices, may become extremely sensitive. With the convergence between different technologies and the growing demand for applications that include location tracking capabilities, privacy concerns pertaining to tracking and profiling activities need to be properly addressed. Many authors have already outlined that there is definitely an issue with the fact that online users may not always be aware that the online profiling and tracking activities are happening in the first place (even if the website privacy policy is open about its practices, many studies have shown that consumers don’t read privacy policies). While this (lack of) consent issue is a serious one, this analysis will instead be focused on other issues: firstly, whether profile data is covered under data protection laws; secondly, whether tracking and profiling activities are legal in accordance with data protection laws such as the Personal Information Protection and Electronic Documents Act (PIPEDA)3; and finally, issues pertaining to the management of profile data, more specifically as they relate to granting access to profile data to individuals and what constitutes a reasonable retention period of the profile data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".