Consumer-Mediated Data Exchange for Research: Current State of US Law, Technology, and Trust
Bibliographic record
Abstract
A compendium of US laws and regulations offers increasingly strong support for the concept that researchers can acquire the electronic health record data that their studies need directly from the study participants using technologies and processes called consumer-mediated data exchange. This data acquisition method is particularly valuable for studies that need complete longitudinal electronic records for all their study participants who individually and collectively receive care from multiple providers in the United States. In such studies, it is logistically infeasible for the researcher to receive necessary data directly from each provider, including providers who may not have the capability, capacity, or interest in supporting research. This paper is a tutorial to inform the researcher who faces these data acquisition challenges about the opportunities offered by consumer-mediated data exchange. It outlines 2 approaches and reviews the current state of provider- and consumer-facing technologies that are necessary to support each approach. For one approach, the technology is developed and estimated to be widely available but could raise trust concerns among research organizations or their institutional review boards because of the current state of US law applicable to consumer-facing technologies. For the other approach, which does not elicit the same trust concerns, the necessary technology is emerging and a pilot is underway. After reading this paper, the researcher who has not been following these developments should have a good understanding of the legal, regulatory, technology, and trust issues surrounding consumer-mediated data exchange for research, with an awareness of what is potentially possible now, what is not possible now, and what could change in the future. The researcher interested in trying consumer-mediated data exchange will also be able to anticipate and respond to an anticipated barrier: the trust concerns that their own organizations could raise.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.283 | 0.359 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.006 | 0.009 |
| Science and technology studies | 0.009 | 0.050 |
| Scholarly communication | 0.034 | 0.054 |
| Open science | 0.008 | 0.022 |
| Research integrity | 0.019 | 0.026 |
| Insufficient payload (model declined to judge) | 0.009 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".