MétaCan
Menu
Back to cohort
Record W4210256061 · doi:10.3390/su14031431

Spatiotemporal Statistical Imbalance: A Long-Term Neglected Defect in UN Comtrade Dataset

2022· article· en· W4210256061 on OpenAlexaboutno aff
Luoming Hu, Changqing Song, Sijing Ye, Peichao Gao

Bibliographic record

VenueSustainability · 2022
Typearticle
Languageen
FieldEconomics, Econometrics and Finance
TopicEconomic and Technological Innovation
Canadian institutionsnot available
FundersBeijing Normal UniversityNational Natural Science Foundation of China
KeywordsCommodityChinaStatistical analysisInternational tradeValue (mathematics)Statistical evidenceCluster analysisWorld tradeOfficial statisticsEconomicsGeographyEconometricsStatisticsMathematicsNull hypothesis

Abstract

fetched live from OpenAlex

The bilateral trade data provided by the United Nations International Trade Statistics Database are some of the most authoritative trade statistics and have been widely used in many research fields. Here, we propose a new form of inconsistency in its records, namely statistical imbalance, which refers to the phenomenon of inequality between the import or export trade value of a commodity category and the total value of all its subcategories. We investigated the frequency and spatial-temporal patterns of the statistical imbalances of 15 reporters (i.e., Australia, Brazil, Canada, China, France, Germany, India, the Netherlands, the Rep. of Korea, the Russian Federation, Switzerland, the United Arab Emirates, the United States of America, and Vietnam) from 1996–2016 and explored their distributional differences in commodity categories with a co-clustering algorithm. The results show that statistical imbalance is widespread with obvious clustering patterns. Trade records related to specific categories such as fossil fuels, pharmaceuticals, machinery, and unspecified commodity categories presented severe statistical imbalances, which may lead to erroneous trade research results. Since statistical imbalance is difficult to detect in studies focusing only on specific commodity categories, we suggested that researchers should prescreen the data for statistical imbalance to ensure the validity of their results.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.005
metaresearch head score (Gemma)0.021
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.018
Threshold uncertainty score0.037

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0050.021
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0060.014
Science and technology studies0.0010.001
Scholarly communication0.0030.002
Open science0.0010.002
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0030.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.024
GPT teacher head0.255
Teacher spread0.231 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations6
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueSustainabilitySame topicEconomic and Technological InnovationFrench-language works237,207