MétaCan
Menu
← Back to cohort
Record W4415975286 · doi:10.1101/2025.11.05.25339575

Commercial or industrial use of mental health data for research: primer and best-practice guidelines from the DATAMIND patient/public Lived Experience Advisory Group

2025· preprint· en· W4415975286 on OpenAlexfundno aff
Linda A. Jones, Naomi Clements-Brod, Jan M. Davies, Marcos DelPozo‐Baños, Andrea Hughes, Matthew H. Iveson, Ann John, Aideen Maguire, Clara Martins de Barros, Andrew M. McIntosh, Michael McTernan, Jan Speechley, Robert Stewart, Jane Taylor, Louise Ting, Rudolf N. Cardinal

Bibliographic record

VenuemedRxiv · 2025
Typepreprint
Languageen
FieldHealth Professions
TopicMental Health and Patient Involvement
Canadian institutionsnot available
FundersQueen's UniversityQueen's University BelfastMedical Research CouncilNHS Health Scotland
KeywordsScope (computer science)Mental healthHealth informaticsFocus groupPublic healthInformaticsData collectionHealth careQualitative property

Abstract

fetched live from OpenAlex

BACKGROUND: Routinely collected health data, such as that held by the United Kingdom (UK) National Health Service/Health and Social Care (collectively "NHS"), has important research uses, but its appropriate use requires public trust and transparency. Commercial/industrial access to routinely collected health data is especially controversial and sensitive for the public, and particular concerns may relate to mental health (MH) data. Existing best-practice MH data science guidelines do not cover commercial uses specifically, but emphasise the importance of patient/public co-development of data science. OBJECTIVES: To develop patient/public-led guidelines for the commercial/industrial use of MH data for research, and to capture relevant background information required by patient/public participants. The focus was on the UK and its constituent nations, but the principles may have wider applicability. METHODS: A patient/public lived experience advisory group (LEAG) was set up within DATAMIND, the Health Data Research UK data hub for MH informatics research development. Initial training and discussion yielded a requirement for definitions and explanations of concepts and processes relating to MH data research, developed iteratively. Subsequently, the LEAG developed guidelines via a qualitative and iterative quasi-Delphi approach. The agreed scope excluded data provided for research with informed consent, data processing arrangements such as companies hosting electronic health records or e-mail systems on the instruction of health services, or compliance with legal minimum requirements. The scope included the use of routinely collected MH data (e.g. NHS data) for research by commercial/industrial organisations without explicit consent, and aspects of MH data collection directly by industry with consent. RESULTS: Alongside the primer in MH data research concepts, the LEAG provide recommendations and best-practice guidelines relating to commercial/industrial research use of MH data, for organisations controlling MH data (such as NHS bodies) and for commercial applicants seeking to use MH data for research. Alongside principles of transparency, patient rights, patient/public involvement in research, stringent governance, and statistical disclosure control, the guidelines recommend a risk-benefit approach to assessing applications for data use, within limits that include avoiding the export of unconsented patient-level data outside NHS-controlled secure data environments, and not providing access to unconsented free-text MH data to commercial applicants. We also provide some recommendations for NHS executive and regulatory bodies, relating to public choice and transparency, clarity of guidance to research-active NHS organisations, and support for de-identification. CONCLUSIONS: Patient/public involvement and understanding is central to MH data research. The primer materials developed here constitute information requested by public advisers prior to considering best practice. The guidelines reflect the views of people with personal or family experience of mental ill health. We hope they are of practical use to the wider MH research community and serve to increase public transparency and trust.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.467
metaresearch head score (Gemma)0.357
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.533
Threshold uncertainty score0.658

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.4670.357
Meta-epidemiology (narrow)0.0020.004
Meta-epidemiology (broad)0.0030.004
Bibliometrics0.0080.009
Science and technology studies0.0090.020
Scholarly communication0.0220.026
Open science0.0140.036
Research integrity0.0310.032
Insufficient payload (model declined to judge)0.0080.009

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.919
GPT teacher head0.618
Teacher spread0.301 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designNot applicable
DomainMethods
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venuemedRxiv→Same topicMental Health and Patient Involvement→French-language works237,207→