On the privacy implications of targeted advertising platforms
Notice bibliographique
Résumé
Online advertising---primarily via online advertising platforms such as Facebook, Google, and Twitter---has rapidly grown to dominate the multi-billion dollar advertising industry. The success of these platforms arises in part from the powerful targeted advertising interfaces these platforms have leveraged their detailed user databases to build, allowing advertisers to target users with ads in a fine-grained manner. This dissertation focuses on two of the most prominent targeting features offered by these ad platforms. The older of these two features allowed advertisers to specify the attributes of the users they wish to target (e.g., their location, relationship status, interests, and income); we call this feature \emph{attribute-based targeting}. The second feature, more recently introduced and potentially more powerful, allows advertisers to specify exactly those users they want to target. This typically works by having the advertiser upload a list of personally identifying information (PII), such as email addresses and phone numbers; the platform then internally matches the PII to users on the platform. This feature, which we call \emph{PII-based targeting}, is particularly popular with advertisers for two key reasons: (i) it allows them to target their existing customers and, (ii) it allows them to target customized lists of users leveraging the myriad sources of personal data that exist today (e.g., voter records, property records, data brokers, criminal records). This dissertation is motivated by \emph{two} high-level privacy concerns arising from these targeting features: \emph{First}, the provision of such targeting features to advertisers (who are typically required to undergo little to no verification) could expose platforms to novel forms of abuse. \emph{Second}, these targeting features rely on detailed data about users; this need for data could incentivize unscrupulous data collection practices by advertising platforms, and by other entities (such as data brokers) that provide data to the advertising ecosystem. Thus, it is essential to understand how these ad platforms work, especially in terms of their data collection and use. \emph{With this motivation, this dissertation posits that the powerful functionalities offered by these fine-grained targeting features open up ad platforms to novel vectors of abuse. Besides, these targeting features raise privacy concerns regarding the data collected to support them. However, at the same time, these very same functionalities provided to advertisers place constraints upon ad platforms; these constraints could potentially be exploited to enforce transparency on ad platforms from the outside.} To this end, this dissertation primarily focuses on Facebook (the largest and most mature of these platforms) and makes \emph{five} key contributions. These contributions collectively surface novel vectors of abuse, propose novel methodologies to enforce transparency on ad platforms, and use these methodologies to discover multiple privacy concerns regarding data collection. All of these contributions only rely on the targeting features provided by ad platforms, and on the size statistics provided by ad platforms to help advertisers plan and optimize their campaigns. In particular, my thesis discovers two novel abuse vectors that abuse these targeting features: \begin{itemize} \item I demonstrate for the first time how size statistics provided in conjunction with PII-based targeting could lead to serious privacy leaks (potentially leaking users' entire PII, besides allowing attackers to de-anonymize visitors to their website en-masse). % I propose a robust fix to such PII-leakage attacks, a stricter variant of which was deployed by Facebook. % %Given the various privacy concerns with targeted advertising platforms, I propose a novel methodologies to remedy the currently limited transparency in three ways (which are the three remaining key contributions): % \item I demonstrate the potential for abusing ad targeting features to target ads in a discriminatory manner (selectively including or excluding users of a particular protected class such as race or gender) across various major advertising platforms (i.e., Facebook, Google, and LinkedIn); in addition, I demonstrate how compositions of individual targeting options could further exacerbate this vector of abuse. % \end{itemize} In addition, my thesis proposes three transparency methodologies that exploit constraints on the delivery of targeted ads, and on the accuracy of size statistics, to bring transparency to three key aspects of ad platforms' data collection: \begin{itemize} % \item I propose a novel methodology to study whether a given potential source of PII is actually used by Facebook to collect PII for PII-based targeting; using the methodology, I demonstrate concerning uses, such as of PII collected for security purposes, and of PII collected without a user's knowledge. % \item I propose a novel methodology to reveal to individual users what data about them in particular is being used for targeted advertising. % \item I propose novel methodologies to audit the extent of data collection, and the accuracy of data used in advertising platforms; I use these methodologies to audit a primary source of data that supports advertising platforms --- data sourced from offline data brokers such as Acxiom and Datalogix --- thereby providing one of the first views of the extent and accuracy of data brokers' data collection (which has traditionally been very opaque). % %problem extend the current discussion in media and literature about discriminatory advertising, showing that we need to focus on features collectively, as compositions of targeting options could exacerbate discriminatory advertising. %In the remainder of this thesis, I will study the potential for a malicious advertiser to covertly run ads in a discriminatory manner, i.e., selectively include or exclude users belonging to a particular sensitive demographic, via various sophisticated attacks. \end{itemize} The results of this dissertation have significantly helped enhance privacy on Facebook's advertising platform and contributed to awareness among the public and regulators about the privacy implications of ad platforms. % In addition, the transparency methodologies I propose in this dissertation alleviate external auditors' sole reliance on the limited transparency mechanisms provided by these ad platforms.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,002 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,001 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».