Bibliographic record
Abstract
This thesis contains three chapters that investigate issues in international trade and innovation. The first chapter examines government-backed technologies promoted in an international standard-setting organization (SSO) and their impacts on global innovation. Standardization ensures compatibility but can steer innovation by locking in specific technologies. Unlike other countries, the Chinese government coordinates its firms to advance specific domestic technologies in international SSOs. I study the impact of this policy on the direction of global innovation in 5G. In the SSO for 5G, firms compete to have their patented technologies adopted as part of the 5G standards. Using a large language model, I build a new database linking SSO technical documents, 5G policy documents published by the Chinese government, and 5G patents. I show that the policy promotes Chinese technologies in areas where China lags behind foreign competitors. If adopted as standards, these lagging technologies become the basis for subsequent 5G innovation across countries. These follow-on patents account for two-fifths of total 5G patents filed worldwide after standardization. In the second chapter, I develop a trade model that incorporates cross-country differences in cost and risk profiles to analyze how risk-averse firms make optimal sourcing decisions considering both cost minimization and risk diversification. Firms are incentivized to minimize expected costs by sourcing from the lowest-cost suppliers, while also reducing profit variance by sourcing additionally from higher-cost alternatives. Calibrating the model to U.S. import data, I show that risk diversification explains 30% of the observed variation in U.S. import shares across countries for intermediate inputs. The third chapter, based on joint work with John Lester, studies technology spillovers from firms performing research and development (R&D) in Canada and their implications for size-based R&D policies. Using panel data covering all firms performing R&D in Canada, we estimate the external return to R&D by size of firm, defining the spillover pool using a measure of technological proximity based on firms' reported expenditure in 147 research fields. We find that spillovers rise with the size of R&D performers, so Canada's policy of subsidizing R&D performed by small firms at a higher rate is not warranted.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.005 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.006 | 0.006 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.016 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".