The U.S. Food and Drug Administration's Mini‐Sentinel program: status and direction
Bibliographic record
Abstract
The Mini-Sentinel is a pilot program that is developing methods, tools, resources, policies, and procedures to facilitate the use of routinely collected electronic healthcare data to perform active surveillance of the safety of marketed medical products, including drugs, biologics, and medical devices. The U.S. Food and Drug Administration (FDA) initiated the program in 2009 as part of its Sentinel Initiative, in response to a Congressional mandate in the FDA Amendments Act of 2007. After two years, Mini-Sentinel includes 31 academic and private organizations. It has developed policies, procedures, and technical specifications for developing and operating a secure distributed data system comprised of separate data sets that conform to a common data model covering enrollment, demographics, encounters, diagnoses, procedures, and ambulatory dispensing of prescription drugs. The distributed data sets currently include administrative and claims data from 2000 to 2011 for over 300 million person-years, 2.4 billion encounters, 38 million inpatient hospitalizations, and 2.9 billion dispensings. Selected laboratory results and vital signs data recorded after 2005 are also available. There is an active data quality assessment and characterization program, and eligibility for medical care and pharmacy benefits is known. Systematic reviews of the literature have assessed the ability of administrative data to identify health outcomes of interest, and procedures have been developed and tested to obtain, abstract, and adjudicate full-text medical records to validate coded diagnoses. Mini-Sentinel has also created a taxonomy of study designs and analytical approaches for many commonly occurring situations, and it is developing new statistical and epidemiologic methods to address certain gaps in analytic capabilities. Assessments are performed by distributing computer programs that are executed locally by each data partner. The system is in active use by FDA, with the majority of assessments performed using customizable, reusable queries (programs). Prospective and retrospective assessments that use customized protocols are conducted as well. To date, several hundred unique programs have been distributed and executed. Current activities include active surveillance of several drugs and vaccines, expansion of the population, enhancement of the common data model to include additional types of data from electronic health records and registries, development of new methodologic capabilities, and assessment of methods to identify and validate additional health outcomes of interest.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | no category Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Not applicable | low |
| gpt | no category Domain: not available · Genre: Commentary About the Canadian research system: no · About a Canadian topic: no | Not applicable | low |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.143 | 0.076 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.004 | 0.007 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.011 | 0.009 |
| Open science | 0.009 | 0.006 |
| Research integrity | 0.011 | 0.011 |
| Insufficient payload (model declined to judge) | 0.015 | 0.008 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".