MétaCan
Menu
Retour à la cohorte
Enregistrement W4248184377 · doi:10.17760/d20394153

On the privacy implications of targeted advertising platforms

2020· dissertation· en· W4248184377 sur OpenAlexaff
Giridhari Venkatadri

Notice bibliographique

Revuenon disponible
Typedissertation
Langueen
DomaineSocial Sciences
ThématiquePrivacy, Security, and Data Protection
Établissements canadiensScience North
Organismes subventionnairesnon disponible
Mots-clésUploadTargeted advertisingComputer scienceOnline advertisingFeature (linguistics)AdvertisingWorld Wide WebPhoneLiberian dollarInternet privacyKey (lock)Contextual advertisingConsumer privacyInformation privacyBusinessThe InternetComputer security

Résumé

récupéré en direct d'OpenAlex

Online advertising---primarily via online advertising platforms such as Facebook, Google, and Twitter---has rapidly grown to dominate the multi-billion dollar advertising industry. The success of these platforms arises in part from the powerful targeted advertising interfaces these platforms have leveraged their detailed user databases to build, allowing advertisers to target users with ads in a fine-grained manner. This dissertation focuses on two of the most prominent targeting features offered by these ad platforms. The older of these two features allowed advertisers to specify the attributes of the users they wish to target (e.g., their location, relationship status, interests, and income); we call this feature \emph{attribute-based targeting}. The second feature, more recently introduced and potentially more powerful, allows advertisers to specify exactly those users they want to target. This typically works by having the advertiser upload a list of personally identifying information (PII), such as email addresses and phone numbers; the platform then internally matches the PII to users on the platform. This feature, which we call \emph{PII-based targeting}, is particularly popular with advertisers for two key reasons: (i) it allows them to target their existing customers and, (ii) it allows them to target customized lists of users leveraging the myriad sources of personal data that exist today (e.g., voter records, property records, data brokers, criminal records). This dissertation is motivated by \emph{two} high-level privacy concerns arising from these targeting features: \emph{First}, the provision of such targeting features to advertisers (who are typically required to undergo little to no verification) could expose platforms to novel forms of abuse. \emph{Second}, these targeting features rely on detailed data about users; this need for data could incentivize unscrupulous data collection practices by advertising platforms, and by other entities (such as data brokers) that provide data to the advertising ecosystem. Thus, it is essential to understand how these ad platforms work, especially in terms of their data collection and use. \emph{With this motivation, this dissertation posits that the powerful functionalities offered by these fine-grained targeting features open up ad platforms to novel vectors of abuse. Besides, these targeting features raise privacy concerns regarding the data collected to support them. However, at the same time, these very same functionalities provided to advertisers place constraints upon ad platforms; these constraints could potentially be exploited to enforce transparency on ad platforms from the outside.} To this end, this dissertation primarily focuses on Facebook (the largest and most mature of these platforms) and makes \emph{five} key contributions. These contributions collectively surface novel vectors of abuse, propose novel methodologies to enforce transparency on ad platforms, and use these methodologies to discover multiple privacy concerns regarding data collection. All of these contributions only rely on the targeting features provided by ad platforms, and on the size statistics provided by ad platforms to help advertisers plan and optimize their campaigns. In particular, my thesis discovers two novel abuse vectors that abuse these targeting features: \begin{itemize} \item I demonstrate for the first time how size statistics provided in conjunction with PII-based targeting could lead to serious privacy leaks (potentially leaking users' entire PII, besides allowing attackers to de-anonymize visitors to their website en-masse). % I propose a robust fix to such PII-leakage attacks, a stricter variant of which was deployed by Facebook. % %Given the various privacy concerns with targeted advertising platforms, I propose a novel methodologies to remedy the currently limited transparency in three ways (which are the three remaining key contributions): % \item I demonstrate the potential for abusing ad targeting features to target ads in a discriminatory manner (selectively including or excluding users of a particular protected class such as race or gender) across various major advertising platforms (i.e., Facebook, Google, and LinkedIn); in addition, I demonstrate how compositions of individual targeting options could further exacerbate this vector of abuse. % \end{itemize} In addition, my thesis proposes three transparency methodologies that exploit constraints on the delivery of targeted ads, and on the accuracy of size statistics, to bring transparency to three key aspects of ad platforms' data collection: \begin{itemize} % \item I propose a novel methodology to study whether a given potential source of PII is actually used by Facebook to collect PII for PII-based targeting; using the methodology, I demonstrate concerning uses, such as of PII collected for security purposes, and of PII collected without a user's knowledge. % \item I propose a novel methodology to reveal to individual users what data about them in particular is being used for targeted advertising. % \item I propose novel methodologies to audit the extent of data collection, and the accuracy of data used in advertising platforms; I use these methodologies to audit a primary source of data that supports advertising platforms --- data sourced from offline data brokers such as Acxiom and Datalogix --- thereby providing one of the first views of the extent and accuracy of data brokers' data collection (which has traditionally been very opaque). % %problem extend the current discussion in media and literature about discriminatory advertising, showing that we need to focus on features collectively, as compositions of targeting options could exacerbate discriminatory advertising. %In the remainder of this thesis, I will study the potential for a malicious advertiser to covertly run ads in a discriminatory manner, i.e., selectively include or exclude users belonging to a particular sensitive demographic, via various sophisticated attacks. \end{itemize} The results of this dissertation have significantly helped enhance privacy on Facebook's advertising platform and contributed to awareness among the public and regulators about the privacy implications of ad platforms. % In addition, the transparency methodologies I propose in this dissertation alleviate external auditors' sole reliance on the limited transparency mechanisms provided by these ad platforms.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,000
score de la tête « metaresearch » (Gemma)0,002
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Théorique ou conceptuel · Signal consensuel: Théorique ou conceptuel
GenreSignal candidat: Empirique · Signal consensuel: aucune
Score de désaccord entre enseignants0,716
Score d'incertitude au seuil0,685

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0000,002
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0010,000
Communication savante0,0000,000
Science ouverte0,0010,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0010,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,033
Tête enseignante GPT0,323
Écart entre enseignants0,289 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeThéorique ou conceptuel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2020
Routes d'admission1
Résumé présentoui

Explorer davantage

Même sujetPrivacy, Security, and Data ProtectionTravaux en français237 207