Commercial or industrial use of mental health data for research: primer and best-practice guidelines from the DATAMIND patient/public Lived Experience Advisory Group
Notice bibliographique
Résumé
BACKGROUND: Routinely collected health data, such as that held by the United Kingdom (UK) National Health Service/Health and Social Care (collectively "NHS"), has important research uses, but its appropriate use requires public trust and transparency. Commercial/industrial access to routinely collected health data is especially controversial and sensitive for the public, and particular concerns may relate to mental health (MH) data. Existing best-practice MH data science guidelines do not cover commercial uses specifically, but emphasise the importance of patient/public co-development of data science. OBJECTIVES: To develop patient/public-led guidelines for the commercial/industrial use of MH data for research, and to capture relevant background information required by patient/public participants. The focus was on the UK and its constituent nations, but the principles may have wider applicability. METHODS: A patient/public lived experience advisory group (LEAG) was set up within DATAMIND, the Health Data Research UK data hub for MH informatics research development. Initial training and discussion yielded a requirement for definitions and explanations of concepts and processes relating to MH data research, developed iteratively. Subsequently, the LEAG developed guidelines via a qualitative and iterative quasi-Delphi approach. The agreed scope excluded data provided for research with informed consent, data processing arrangements such as companies hosting electronic health records or e-mail systems on the instruction of health services, or compliance with legal minimum requirements. The scope included the use of routinely collected MH data (e.g. NHS data) for research by commercial/industrial organisations without explicit consent, and aspects of MH data collection directly by industry with consent. RESULTS: Alongside the primer in MH data research concepts, the LEAG provide recommendations and best-practice guidelines relating to commercial/industrial research use of MH data, for organisations controlling MH data (such as NHS bodies) and for commercial applicants seeking to use MH data for research. Alongside principles of transparency, patient rights, patient/public involvement in research, stringent governance, and statistical disclosure control, the guidelines recommend a risk-benefit approach to assessing applications for data use, within limits that include avoiding the export of unconsented patient-level data outside NHS-controlled secure data environments, and not providing access to unconsented free-text MH data to commercial applicants. We also provide some recommendations for NHS executive and regulatory bodies, relating to public choice and transparency, clarity of guidance to research-active NHS organisations, and support for de-identification. CONCLUSIONS: Patient/public involvement and understanding is central to MH data research. The primer materials developed here constitute information requested by public advisers prior to considering best practice. The guidelines reflect the views of people with personal or family experience of mental ill health. We hope they are of practical use to the wider MH research community and serve to increase public transparency and trust.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,467 | 0,357 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,004 |
| Méta-épidémiologie (sens large) | 0,003 | 0,004 |
| Bibliométrie | 0,008 | 0,009 |
| Études des sciences et des technologies | 0,009 | 0,020 |
| Communication savante | 0,022 | 0,026 |
| Science ouverte | 0,014 | 0,036 |
| Intégrité de la recherche | 0,031 | 0,032 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,008 | 0,009 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; l’étiquette directe de Gemma et le classifieur distillé Codex s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».