Estimating Trade Flows: Trading Partners and Trading Volumes<sup>*</sup>
Bibliographic record
Abstract
We develop a simple model of international trade with heterogeneous firms that is consistent with a number of stylized features of the data. In particular, the model predicts positive as well as zero trade flows across pairs of countries, and it allows the number of exporting firms to vary across destination countries. As a result, the impact of trade frictions on trade flows can be decomposed into the intensive and extensive margins, where the former refers to the trade volume per exporter and the latter refers to the number of exporters. This model yields a generalized gravity equation that accounts for the self-selection of firms into export markets and their impact on trade volumes. We then develop a two-stage estimation procedure that uses an equation for selection into trade partners in the first stage and a trade flow equation in the second. We implement this procedure parametrically, semiparametrically, and nonparametrically, showing that in all three cases the estimated effects of trade frictions are similar. Importantly, our method provides estimates of the intensive and extensive margins of trade. We show that traditional estimates are biased and that most of the bias is due not to selection but rather due to the omission of the extensive margin. Moreover, the effect of the number of exporting firms varies across country pairs according to their characteristics. This variation is large and particularly so for trade between developed and less developed countries and between pairs of less developed countries.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.015 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.008 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".