Bibliographic record
Abstract
Online advertising---primarily via online advertising platforms such as Facebook, Google, and Twitter---has rapidly grown to dominate the multi-billion dollar advertising industry. The success of these platforms arises in part from the powerful targeted advertising interfaces these platforms have leveraged their detailed user databases to build, allowing advertisers to target users with ads in a fine-grained manner. This dissertation focuses on two of the most prominent targeting features offered by these ad platforms. The older of these two features allowed advertisers to specify the attributes of the users they wish to target (e.g., their location, relationship status, interests, and income); we call this feature \emph{attribute-based targeting}. The second feature, more recently introduced and potentially more powerful, allows advertisers to specify exactly those users they want to target. This typically works by having the advertiser upload a list of personally identifying information (PII), such as email addresses and phone numbers; the platform then internally matches the PII to users on the platform. This feature, which we call \emph{PII-based targeting}, is particularly popular with advertisers for two key reasons: (i) it allows them to target their existing customers and, (ii) it allows them to target customized lists of users leveraging the myriad sources of personal data that exist today (e.g., voter records, property records, data brokers, criminal records). This dissertation is motivated by \emph{two} high-level privacy concerns arising from these targeting features: \emph{First}, the provision of such targeting features to advertisers (who are typically required to undergo little to no verification) could expose platforms to novel forms of abuse. \emph{Second}, these targeting features rely on detailed data about users; this need for data could incentivize unscrupulous data collection practices by advertising platforms, and by other entities (such as data brokers) that provide data to the advertising ecosystem. Thus, it is essential to understand how these ad platforms work, especially in terms of their data collection and use. \emph{With this motivation, this dissertation posits that the powerful functionalities offered by these fine-grained targeting features open up ad platforms to novel vectors of abuse. Besides, these targeting features raise privacy concerns regarding the data collected to support them. However, at the same time, these very same functionalities provided to advertisers place constraints upon ad platforms; these constraints could potentially be exploited to enforce transparency on ad platforms from the outside.} To this end, this dissertation primarily focuses on Facebook (the largest and most mature of these platforms) and makes \emph{five} key contributions. These contributions collectively surface novel vectors of abuse, propose novel methodologies to enforce transparency on ad platforms, and use these methodologies to discover multiple privacy concerns regarding data collection. All of these contributions only rely on the targeting features provided by ad platforms, and on the size statistics provided by ad platforms to help advertisers plan and optimize their campaigns. In particular, my thesis discovers two novel abuse vectors that abuse these targeting features: \begin{itemize} \item I demonstrate for the first time how size statistics provided in conjunction with PII-based targeting could lead to serious privacy leaks (potentially leaking users' entire PII, besides allowing attackers to de-anonymize visitors to their website en-masse). % I propose a robust fix to such PII-leakage attacks, a stricter variant of which was deployed by Facebook. % %Given the various privacy concerns with targeted advertising platforms, I propose a novel methodologies to remedy the currently limited transparency in three ways (which are the three remaining key contributions): % \item I demonstrate the potential for abusing ad targeting features to target ads in a discriminatory manner (selectively including or excluding users of a particular protected class such as race or gender) across various major advertising platforms (i.e., Facebook, Google, and LinkedIn); in addition, I demonstrate how compositions of individual targeting options could further exacerbate this vector of abuse. % \end{itemize} In addition, my thesis proposes three transparency methodologies that exploit constraints on the delivery of targeted ads, and on the accuracy of size statistics, to bring transparency to three key aspects of ad platforms' data collection: \begin{itemize} % \item I propose a novel methodology to study whether a given potential source of PII is actually used by Facebook to collect PII for PII-based targeting; using the methodology, I demonstrate concerning uses, such as of PII collected for security purposes, and of PII collected without a user's knowledge. % \item I propose a novel methodology to reveal to individual users what data about them in particular is being used for targeted advertising. % \item I propose novel methodologies to audit the extent of data collection, and the accuracy of data used in advertising platforms; I use these methodologies to audit a primary source of data that supports advertising platforms --- data sourced from offline data brokers such as Acxiom and Datalogix --- thereby providing one of the first views of the extent and accuracy of data brokers' data collection (which has traditionally been very opaque). % %problem extend the current discussion in media and literature about discriminatory advertising, showing that we need to focus on features collectively, as compositions of targeting options could exacerbate discriminatory advertising. %In the remainder of this thesis, I will study the potential for a malicious advertiser to covertly run ads in a discriminatory manner, i.e., selectively include or exclude users belonging to a particular sensitive demographic, via various sophisticated attacks. \end{itemize} The results of this dissertation have significantly helped enhance privacy on Facebook's advertising platform and contributed to awareness among the public and regulators about the privacy implications of ad platforms. % In addition, the transparency methodologies I propose in this dissertation alleviate external auditors' sole reliance on the limited transparency mechanisms provided by these ad platforms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".