Quality evaluation of field safety notices of medical devices, across EU and non-EU countries
Bibliographic record
Abstract
Abstract Introduction The EU Medical Device Regulation 2017/745 (MDR; article 87) requires device manufacturers to report to the relevant authorities any serious incident and any corrective action undertaken, through Field Safety Notices (FSN) whose content must be consistent in all Member States. Before MDR, due to the lack of harmonized global standards for reporting, FSNs were published independently by each country with different formats, styles, nomenclatures and languages. This heterogeneity makes it challenging to use such historical data for trend analysis in post-market surveillance (PMS). Purpose 1) To assess the quality of the FSNs issued by EU or non-EU national authorities for analysing historical trends. 2) To group countries based on such quality, and test for differences between EU and non-EU countries. Methods As part of the CORE-MD project coordinated by the ESC, all FSN information publicly available as HTML text (excluding linked PDFs) was retrieved automatically from national authorities’ websites, using the CORE-MD PMS tool [1]. For each FSN, a score was assigned to each field of interest (1 or 2, according to importance for building trends) (see Figure 1), and a Percentage Score (PS) computed as the % of the ratio between the scores’ sum and the maximum possible score (10). Based on PS, FSNs were categorized as Excellent (≥90), Very Good (80≤PS<90), Good (70≤PS<80), Medium (60≤PS<70), or Unqualified (<60). For each country, the distribution of FSN categories was computed. Hierarchical Agglomerative Clustering (HAC) was used to group countries based on their similarities, and differences between EU and non-EU countries were assessed by Mann-Whitney U test. Results 126,405 FSNs published before 31/12/2022 were retrieved and analyzed. The % of FSN in which each field was available is shown by country in Figure 1. Manufacturer and Device were clearly described, but more detailed device information was often absent. The Description, with the reason for publishing the FSN, and the Device Category, allowing nomenclature categorization, were provided only by a few countries. The % of FSNs in each quality category is reported in Figure 2: only Italy provided some FSNs (31%) classified as Excellent. The HAC analysis identified three clusters: "high-quality FSNs" (Czechia, Denmark, Italy, Sweden, the Netherlands, and the USA), "moderate-quality FSNs" (Australia, Canada, Greece, Ireland, Latvia, Slovenia, Spain, and the UK), and "low-quality FSNs" (Croatia, Estonia, France, Germany, Poland, and Portugal). There was no difference between EU and non-EU countries. Conclusions A significant discrepancy was observed in the quality of FSNs retrieved from different countries, highlighting the difficulty in using such data to analyse trends, for example in reports of cardiovascular devices. This study reinforces the need for a global minimum reporting standard, to facilitate more effective PMS.Figure 1Figure 2
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.073 | 0.182 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.015 | 0.014 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".