Global Soybean Research Dynamic Analysis Based on SCI-EXPANDED Database
Bibliographic record
Abstract
We analyzed the amount of soybean-related articles and the citation of research countries,institutions,authors,and journals based on SCI-EXPANDED database of Web of Science. Literature co-citation maps associated with soybean research and knowledge were generated by CiteSpace III information visualization software,and it directly demonstrated the knowledge base and research forefront of soybean. The results showed that a total of 17 576 soybean research articles from 122 countries were published,with the 36 190 authors belonging to 5 879 organizations. The number of articles increased by year in general, and the top five countries by article number were the USA,Brazil,China,Japan,and South Korea. In terms of article quality, the United States Department of Agriculture,Iowa State University,and University of Illinois excelled to other organizations. Hartman G L,Shoemaker R C,Boerma H R,and Nelson R L from USA,and Cober E R from Canada were the top five productive authors for soybean research in the world. Articles were most frequently published in Crop Science,Journal of Agricultural and Food Chemistry,Journal of the American Oil Chemists Society,Agronomy Journal and Pesquisa Agropecuaria Brasileira,focusing on the categories of plant sciences,agronomy,food science technology,agriculture-multidisciplinary,chemistry-applied,and biochemistry molecular biology. The research forefronts covered the areas of soybean molecular biology,plant protection,genetics and breeding,processing and quality,transformation technology,soybean meal,and biodiesel. In general,China has ranked high in soybean research scale and levels in terms of publication number,but the number of high-impacted articles was insufficient. It was proposed that more efforts should be put on soybean research in China to improve its international influence.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.012 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.101 | 0.115 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.011 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".