Authors' reply: Are low‐ and middle‐income countries achieving the Lancet commission global benchmark for surgical volumes? A systematic review
Bibliographic record
Abstract
We thank Davis et al.1 for their interest and insightful feedback on our work about surgical volumes in low- and middle-income countries (LMICs).2 We appreciate Davis et al. for highlighting the importance of including nonacademic publications in the search strategy. The considerable increase in the number of countries reporting from 12 to 43 and surgeries from 877 to 1367 from these countries is expected due to the inclusion of data from ministry of health reports, national statistical databases, and gray literature from these countries. Our systematic review aimed to report and analyze surgical volume in LMICs, assess proxy indicators, and document the limitations and barriers to surgical volume data collection based on peer-reviewed scientific research.3 While we acknowledge that MOHs and national agencies often shoulder the responsibility of data collection in LMICs, there are major challenges in using this data for benchmarking or wider comparisons. Firstly, the bias in reporting, overreporting in particular, is a major challenge with this self-reported data. Secondly, not all countries report surgical volumes similarly, and denominators, methods and criteria for inclusion, and definitions of surgical procedures may vary considerably, affecting data comparability.4 Standardizing these diverse sources to ensure quality and reliability comparable to peer-reviewed publications is a challenge. Despite excluding gray literature, the heterogeneity in the collected data is a major barrier and challenge reported by our systematic review. Our study, hence, focused on peer-reviewed literature for methodological consistency and reliability. Also, the triangulation of data sources enhances understanding provided by self-reported data from government agencies. The discrepancy between academic publications and government reports underscores the need for harmonized data collection and reporting mechanisms. Standardizing data templates and incorporating them into existing government systems would pave the way for uniform definitions and reporting and would enhance the quality and availability of surgical data. This will also facilitate better monitoring and evaluation of progress in this important surgical indicator.5 In conclusion, we commend Davis et al. for their comprehensive review and strategy to collaborate with MOHs to improve data collection and reporting. Priti Patil: Conceptualization; data curation; formal analysis; writing – original draft; writing – review & editing. Priyansh Nathani: Conceptualization; data curation; formal analysis; writing – original draft; writing – review & editing. Juul M. Bakker: Conceptualization; formal analysis; writing – original draft; writing – review & editing. Alex J. van Duinen: Conceptualization; formal analysis; writing – review & editing. Pranav Bhushan: Writing – review & editing. Minal Shukla: Writing – review & editing. Samir Chalise: Writing – review & editing. Nobhojit Roy: Conceptualization; writing – review & editing. Anita Gadgil: Conceptualization; writing – review & editing. This research did not receive any external funding. The authors declare that they have no conflicts of interest.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.010 | 0.002 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".