Open Science, Data, and Methodologies: Lessons learned from the NIHR-RESPIRE Network in Asia
Bibliographic record
Abstract
NIHR-RESPIRE, a Global Health Research Unit funded by NIHR, is committed to advancing respiratory health research in Asia. We prioritise Open Science, Data, and Methodologies to maximise research data utility securely, sharing the lessons encountered across seven LMIC partner countries (Bangladesh, Bhutan, India, Indonesia, Malaysia, Pakistan, and Sri Lanka). Our strategic shift from traditional data sharing to LMIC-tailored Open Science practices ensures data privacy and security. This includes refining Data Management Plans, metadata standards, and mandating FAIR Data sharing, providing methodological support, and developing Open Science Policy Guidelines. We advocate for the adoption of open science principles to maximise secure data use and value with a focus on FAIR data. We also provide aid to partners in enhancing their data-related skills, hosting regular meetings, and establishing internal data monitoring structures to bolster cross-cutting activities within RESPIRE. Through capacity building, we have enabled high-quality respiratory health research using Open Science principles, enhancing data sharing efficiency, research visibility, and ultimately respiratory health outcomes in Asia and beyond. Our experience underscores the following lessons: Flexibility in data sharing, tailored to LMIC researchers' needs, is essential; Training and support to enhance knowledge of methodologies and dispel misconceptions are key to successful data stewardship; Appointing a focal person for structured anonymised data sharing and supporting the internal Data Monitoring Committee are critical. We recognise Open Science's potential to foster innovation, collaboration, and knowledge sharing in respiratory health research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.303 | 0.162 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.003 | 0.005 |
| Science and technology studies | 0.007 | 0.028 |
| Scholarly communication | 0.030 | 0.039 |
| Open science | 0.007 | 0.032 |
| Research integrity | 0.007 | 0.023 |
| Insufficient payload (model declined to judge) | 0.004 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".