OpenAIRE Research Graph: Dump of funded products
Bibliographic record
Abstract
This dataset contains the metadata records about research products (research literature, data, software, other types of research products) with funding information available in the OpenAIRE Research Graph produced on 13 March 2022.<br> Records are grouped by funder in a dedicated archive file (<funder acronym>.tar). Funder acronym Funder name AKA Academy of Finland ANR French National Research Agency ARC Australian Research Council CHIST-ERA CHIST-ERA CIHR Canadian Institute of Health Research EC_FP7 European Commission FP7 projects EC_H2020 European Commission H2020 projects FCT Fundação para a Ciência e a Tecnologia, I.P. FWF Austrian Science Fund HRZZ Croatian Science Foundation (CSF) MZOS Ministry of Science, Education and Sports of the Republic of Croatia (MSES) MESTD Ministry of Education, Science and Technological Development of Republic of Serbia NHMRC National Health and Medical Research Council NIH National Institutes of Health (US) NSERC Natural Sciences and Engineering Research Council of Canada NSF National Science Foundation (US) NWO Netherlands Organisation for Scientific Research SFI Science Foundation Ireland SNSF Swiss National Science Foundation SSHRC Social Sciences and Humanities Research Council TARA Tara Expeditions Foundation TUBITAK Türkiye Bilimsel ve Teknolojik Araştırma Kurumu UKRI UK Research and Innovation WT Wellcome Trust Each tar archive contains gzip files, each with one json record per line. Json records are compliant with the schema available at https://doi.org/10.5281/zenodo.4723499. You can also search and browse this dataset (and more) in the OpenAIRE EXPLORE portal and via the OpenAIRE API.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.003 | 0.007 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.033 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".