Learning While Bidding in Real Time Auctions with Multiple Item Types and Unknown Price Distribution
Bibliographic record
Abstract
A Real-Time Bidding (RTB) network is a real-time auction market, primarily used for \nadvertising space sales. Within this environment, clients participate by bidding on preferred \nitems and subsequently purchasing them upon winning. This thesis addresses the \nproblem of optimal real-time bidding within a second-price Vickrey auction setting, where \nthe distribution of prices is unknown. Our focus centers on second-price auction mechanisms, \nwhich offers unique properties that enable the development of compelling algorithms. \nWe introduce the concept of a demand-side platform (DSP), acting as an intermediary representing \nclients in the auction market. With no prior knowledge of typical prices, the DSP \nmust determine optimal bidding strategies for each item and distribute won items among \nclients to fulfill their contracts while minimizing expenses. When the distribution of the \nprices of items is known, this optimal bidding problem can be solved by classic convex \noptimization algorithms such as ADMM. However, market properties may vary over time, \nand access to competitor behavior or bidding information is limited. Consequently, the \nDSP must continually update its information about the price distribution, while adapting \nbidding estimations in real-time. Our primary contribution lies in devising efficient \nonline optimization algorithms that accurately find the optimal bids. To tackle this, we \nemploy tools from convex optimization analysis, including duality, along with stochastic \noptimization algorithms, notably stochastic approximation. Moreover, techniques such as \nprojection and penalty term methods are utilized to enhance algorithm performance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.013 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".