Regulatory Issues in Electronic Health Records for Adolescent HIV Research: Strategies and Lessons Learned
Bibliographic record
Abstract
BACKGROUND: Electronic health records (EHRs) are a cost-effective approach to provide the necessary foundations for clinical trial research. The ability to use EHRs in real-world clinical settings allows for pragmatic approaches to intervention studies with the emerging adult HIV population within these settings; however, the regulatory components related to the use of EHR data in multisite clinical trials poses unique challenges that researchers may find themselves unprepared to address, which may result in delays in study implementation and adversely impact study timelines, and risk noncompliance with established guidance. OBJECTIVE: As part of the larger Adolescent Trials Network (ATN) for HIV/AIDS Interventions Protocol 162b (ATN 162b) study that evaluated clinical-level outcomes of an intervention including HIV treatment and pre-exposure prophylaxis services to improve retention within the emerging adult HIV population, the objective of this study is to highlight the regulatory process and challenges in the implementation of a multisite pragmatic trial using EHRs to assist future researchers conducting similar studies in navigating the often time-consuming regulatory process and ensure compliance with adherence to study timelines and compliance with institutional and sponsor guidelines. METHODS: Eight sites were engaged in research activities, with 4 sites selected from participant recruitment venues as part of the ATN, who participated in the intervention and data extraction activities, and an additional 4 sites were engaged in data management and analysis. The ATN 162b protocol team worked with site personnel to establish the necessary regulatory infrastructure to collect EHR data to evaluate retention in care and viral suppression, as well as para-data on the intervention component to assess the feasibility and acceptability of the mobile health intervention. Methods to develop this infrastructure included site-specific training activities and the development of both institutional reliance and data use agreements. RESULTS: Due to variations in site-specific activities, and the associated regulatory implications, the study team used a phased approach with the data extraction sites as phase 1 and intervention sites as phase 2. This phased approach was intended to address the unique regulatory needs of all participating sites to ensure that all sites were properly onboarded and all regulatory components were in place. Across all sites, the regulatory process spanned 6 months for the 4 data extraction and intervention sites, and up to 10 months for the data management and analysis sites. CONCLUSIONS: The process for engaging in multisite clinical trial studies using EHR data is a multistep, collaborative effort that requires proper advanced planning from the proposal stage to adequately implement the necessary training and infrastructure. Planning, training, and understanding the various regulatory aspects, including the necessity of data use agreements, reliance agreements, external institutional review board review, and engagement with clinical sites, are foremost considerations to ensure successful implementation and adherence to pragmatic trial timelines and outcomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.315 | 0.408 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.004 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.008 | 0.017 |
| Scholarly communication | 0.029 | 0.039 |
| Open science | 0.009 | 0.013 |
| Research integrity | 0.018 | 0.027 |
| Insufficient payload (model declined to judge) | 0.007 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".