Bibliographic record
Abstract
Web Single Sign-On (SSO) login is a popular alternative to password-based login in current authentication systems.SSO services enable users to use accounts registered with identity providers (IdPs) such as Google and Facebook to login on multiple relying party (RP) websites.Common web SSO deployments are based on the OAuth 2.0 authorization standard which enables RPs to both authenticate users and access a subset of a user's personal information from an IdP.This thesis pursues three goals related to user privacy in OAuth-based web SSO implementations.First, we build OAuthScope, a tool that extracts OAuth protocol data from RP sites.We use it to conduct an empirical investigation of privacy implications for users of OAuth implementations in RP websites most visited by users across five countries.We categorize user data made available by four IdPs (Google, Facebook, Apple, and LinkedIn) and evaluate the types of user data accessed by RPs through these IdPs.Our results reveal considerable variations in the categories and amounts of user data accessed by RPs, including differences across site versions in different countries.Second, to improve the transparency of user data accessed by RPs, we design and implement SSOPrivateEye (SPEye), a browser extension tool to inform users about the privacy consequences of choosing SSO login options.SPEye extracts information about permission requests made by RPs to enable users to compare SSO options before making a login choice.Third, we conduct a user study to identify factors that influence participants when choosing from SSO and non-SSO login options.We compare login decisions made by participants before and after viewing comparative information on the user data accessed by RPs through different SSO choices.We find that usability preferences and inertia influence a majority of login decisions when presented with a list of SSO and non-SSO choices, while privacy-related reasons were most common after participants viewed the user data requested through each SSO choice.Through these three goals, we highlight and tackle SSO privacy issues affecting users, and provide insights on further improving privacy in OAuth systems.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.071 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.004 | 0.005 |
| Scholarly communication | 0.010 | 0.016 |
| Open science | 0.001 | 0.005 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".