MétaCan
Menu
Back to cohort
Record W6912894846 · doi:10.5683/sp3/0kr09p

Quantifying industry spending on promotional events using Open Payments data: Event classification script

2024· dataset· en· W6912894846 on OpenAlexaboutno aff

Bibliographic record

VenueBorealis · 2024
Typedataset
Languageen
Field
Topic
Canadian institutionsnot available
Fundersnot available
KeywordsPaymentCertificationConsistency (knowledge bases)Event (particle physics)Quarter (Canadian coin)DownloadIdentifierPayment card

Abstract

fetched live from OpenAlex

We conducted a cross-sectional study of the publicly available 2022 Open Payments data to characterize and quantify sponsored events (available for download at: https://www.cms.gov/priorities/key-initiatives/open-payments/data/dataset-downloads). Data sources We downloaded the 2022 dataset ZIP files from the Open Payments website on June 30th, 2023. We included all records for nurse practitioners, clinical nurse specialists, certified registered nurse anesthetists, and certified nurse-midwives (hereafter advanced practiced registered nurses (APRNs)); and allopathic and osteopathic physicians (hereafter, ‘physicians’). To ensure consistency in provider classification, we linked Payments data to the National Plan and Provider Enumeration System data (June 2023) by National Provider Identifier (NPI) and the National Uniform Claim Committee (NUCC) and excluded individuals with an ambiguous provider type. Event-centric analysis of Open Payments records: Creating an event typology We included only payments classified as “food and beverage” to reliably identify distinct sponsored events. We reasoned that food and beverage would be consumed on the same day in the same place, thus assumed that records for food and beverage associated with the same event would share the date of payment and location. We also assumed that the reported value of a food and beverage payment is the total cost of the hospitality divided by the number of attendees, thus grouped payment records with the same amount, rounded to the nearest dollar. Inferring which Open Payment records relate to the same sponsored event requires analytic decisions regarding the selection and representation of variables that define an event. To understand the impact of these choices, we undertook a sensitivity analysis to explore alternative ways to group Open Payments records for food and beverage, to determine how combination of variables, including date (specific date or within the same calendar week), amount (rounded to nearest dollar), and recipient’s state, affected the identification of sponsored events in the Open Payments data set. We chose to define a sponsored event as a cluster of three or more individual payment records for food and beverage (nature of payment) with the following matching Open Payments record variables: • Submitting applicable manufacturer (name) • Product category or therapeutic area • Name of drug or biological or device or medical supply • Recipient state • Total amount of payment (USD, rounded to nearest dollar) • Date of payment (exact) After examining the distribution of the data, we classified events in terms of size (≥20 attendees as “large” and 3-<20 as “small”) and amount per person. We categorized events <$10 as “coffee”, $10-<$30 as “lunch”, $30-<$150 as “dinner”, and ≥$150 as “banquet”.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.003
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Scholarly communication, Research integrity, Insufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.017
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0030.000
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0010.001
Science and technology studies0.0000.000
Scholarly communication0.0010.001
Open science0.0050.004
Research integrity0.0010.003
Insufficient payload (model declined to judge)0.0000.005

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.387
GPT teacher head0.451
Teacher spread0.064 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueBorealisFrench-language works237,207