Improving Accessibility and Usability of Clinical Data for the Swiss Personalized Health Network: Development and Usability Study (Preprint)
Bibliographic record
Abstract
Background: In large-scale research initiatives such as the Swiss Personalized Health Network (SPHN), ensuring interoperability and ease of use across diverse clinical datasets is challenging due to a lack of standardization and semantics. The approach taken by the SPHN with the creation of a reference common dataset integrating various data structures and standards creates complexities for researchers aiming to explore and find specific clinical concepts to assess project feasibility. Semantic enrichment and exploration through SNOMED CT offer a potential solution by enabling structured queries that could simplify data discoverability and enhance dataset usability across Switzerland. Objective: This study evaluates how a semantic layer can improve the exploration of the SPHN dataset by leveraging SNOMED CT's hierarchical terminology. To validate this hypothesis, we developed the Smart SNOMED Search for SPHN (S4) tool, which leverages a semantic enrichment of the dataset, facilitating semantic searches using the Expression Constraint Language of SNOMED CT. Methods: The SPHN dataset underwent semantic enrichment, where concepts and attributes not already represented were systematically mapped to SNOMED CT codes and associated value sets. An additional 717 meaning bindings and 232 value sets were created. The S4 tool was designed to enable Expression Constraint Language-based queries, allowing users to retrieve relevant SPHN concepts and value sets effectively. We tested the tool using a validation dataset representing commonly encountered clinical data warehouse elements and evaluated its precision, recall, and F1-scores. Results: The S4 tool demonstrated high accuracy, with an overall precision of 95.3%, recall of 97.5%, and an F1-score of 96.4%, indicating effective retrieval and alignment of SPHN concepts with SNOMED CT codes. The enrichment also highlighted gaps within the SPHN dataset, such as a lack of representation and the inability to use postcoordination, enabling enhanced semantic connections that further supported data discoverability. Conclusions: The S4 tool validates the hypothesis that semantic representation enhances data explorability within large frameworks such as the SPHN, making it more accessible for research. While effective, future work could address limitations by refining search precision and improving accessibility for users less familiar with SNOMED CT, thereby supporting SPHN's mission to facilitate personalized health care research through enhanced data interoperability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.030 | 0.094 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".