Best practice considerations on the assessment of robotic assisted surgical systems: results from an international consensus expert panel
Bibliographic record
Abstract
BACKGROUND: Health technology assessments (HTAs) of robotic assisted surgery (RAS) face several challenges in assessing the value of robotic surgical platforms. As a result of using different assessment methods, previous HTAs have reached different conclusions when evaluating RAS. While the number of available systems and surgical procedures is rapidly growing, existing frameworks for assessing MedTech provide a starting point, but specific considerations are needed for HTAs of RAS to ensure consistent results. This work aimed to discuss different approaches and produce guidance on evaluating RAS. METHODS: A consensus conference research methodology was adopted. A panel of 14 experts was assembled with international experience and representing relevant stakeholders: clinicians, health economists, HTA practitioners, policy makers, and industry. A review of previous HTAs was performed and seven key themes were extracted from the literature for consideration. Over five meetings, the panel discussed the key themes and formulated consensus statements. RESULTS: A total of ninety-eight previous HTAs were identified from twenty-five total countries. The seven key themes were evidence inclusion and exclusion, patient- and clinician-reported outcomes, the learning curve, allocation of costs, appropriate time horizons, economic analysis methods, and robotic ecosystem/wider benefits. CONCLUSIONS: Robotic surgical platforms are tools, not therapies. Their value varies according to context and should be considered across therapeutic areas and stakeholders. The principles set out in this paper should help HTA bodies at all levels to evaluate RAS. This work may serve as a case study for rapidly developing areas in MedTech that require particular consideration for HTAs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".