Exploring the Role of Terminology in SNOMED CT Definitions: Challenges and Solutions
Notice bibliographique
Résumé
This work aims to show how insights from Terminology Science can foster interoperability and more effective knowledge representation, organisation, and sharing in the biomedical field. Definitions are one of the key forms of concept representation in Terminology and allow a given concept to be identified within a concept system while also enabling its differentiation from other related concepts [1]. They are also important to ensure conceptual and linguistic consistency of specialised terminology. Furthermore, definitions are an essential tool for those who need to access (or work on) a biomedical terminological resource but are not domain experts, such as translators (and patients). However, producing accurate and understandable definitions in natural language (i.e. text format) is a difficult and time-consuming task. Terminology can provide fundamental support in this process. SNOMED CT concepts are defined in three different ways [2]: a) the Fully Specified Name (FSN); b) a formal concept definition; and c) textual definitions. These three types of definitions can be useful for the work carried out by the interdisciplinary teams responsible for translating SNOMED CT content into the various languages of the SNOMED International Member Countries. However, professional translators without medical expertise will have greater difficulty understanding all SNOMED CT concepts starting from the FSN and/or their formal relationships. Moreover, the formal representation of concepts may contain errors which can only be detected by translators with sufficient domain knowledge. Therefore, natural language definitions and textual resources that provide information about the contextual use of a term are key additional tools. However, for the time being, textual definitions in SNOMED CT are optional and only provided for a limited number of concepts, when there is a requirement for additional detail. Having a language-independent, logically consistent, and interoperable conceptual foundation can help optimise the drafting process without jeopardising the added value supplied by linguistic diversity. Anchored in a double-dimensional approach to Terminology [3], which considers, within a given domain, not only the systematic and semi-automatic processing and analysis of specialised texts and their terms but also a (semi)formal representation of the core concepts and their respective organisation into systems, we propose a methodology for textual definition drafting based on the formal definitions in SNOMED CT. One of the examples analysed in the case study refers to |Single-port laparoscopic cholecystectomy (procedure)|, a primitive concept without a natural language definition proposal. To support the drafting of the latter, the formal definition was analysed, as well as its alignment with the existing FSN. However, it became clear that from a terminological standpoint, the concept |Single-port laparoscopic cholecystectomy (procedure)| currently has the same formal definition as its ‘parent’ concept |Laparoscopic cholecystectomy (procedure)| - a fully defined concept -, even though they are, in fact, two different concepts. Furthermore, the delimiting characteristic of the concept under analysis is not represented in the formal definition, despite its presence in the FSN (“single-port”). The proposal to incorporate the essential characteristic of |Single-port laparoscopic cholecystectomy (procedure)| into the formal definition, thereby enabling the subsequent textual definition to reflect that, is based on EndoTerm [4], which combines an Aristotelian-based, specific difference approach in the main hierarchies with a customised set of non-hierarchical concept relations, mostly based on SNOMED CT and UMLS, as well as with a categorial structure, to prevent logical errors. The concept systems within EndoTerm were encoded as a formal ontology, thereby enabling interoperability. An example is provided in Figure 1 for the micro-concept system of |Laparoendoscopic single-site total hysterectomy (procedure)|, a surgical procedure with the same surgical approach as |Single-port laparoscopic cholecystectomy (procedure)|. After the formal definition has been stabilised, a definitional template-like structure was created to optimise the textual definition drafting process and enhance consistency. The notion of definitional templates has been developed in terminology work within Frame-based Terminology [5]. In this work, this definitional template approach is taken further via the incorporation of ontologies and the use of formal SNOMED CT definitions. The definitional template in question can be seen in Figure 2 and based on it, a natural language definition proposal for |Laparoendoscopic single-site total hysterectomy (procedure)| would be “Procedure consisting of the excision of the uterus and cervix using a laparoscope to access the site via a minimally invasive surgical approach with a single skin incision”. Following this approach, only a minor, indispensable change would be required for |Single-port laparoscopic cholecystectomy (procedure)|, namely the body structure being excised. While ideally, the three concept definitions in SNOMED CT should be aligned, in order to facilitate the automation of the textual definition drafting process, this is not always the case, as the analysis of other examples has shown. Given the growing trend in incorporating textual definitions into biomedical terminological resources, as well as the importance that such natural language definitions have for translators and, potentially, for other user groups (namely patients), it is believed that Terminology – with its double-dimensional nature – and ontologies can bring added value from a methodological standpoint, providing consistency to both formal and linguistic aspects of the SNOMED CT definitions and contributing to a more successful alignment between them. References: [1] ISO. ISO 1087:2019. Terminology work and terminology science — Vocabulary, Geneva, 2019. [2] SNOMED International. SNOMED CT Editorial Guide, 2022. URL: https://confluence.ihtsdotools.org/display/DOCEG [3] Santos, C. and Costa, R. “Domain Specificity: Semasiological and Onomasiological Knowledge Representation”. Handbook of Terminology. Ed. Hendrik J. Kockaert and Frieda Steurs. Amsterdam / Philadelphia: John Benjamins, 2015: 153-179. [4] Carvalho, S. A terminological approach to knowledge organization within the scope of endometriosis: the EndoTerm project. PhD, Univ. NOVA de Lisboa - FCSH / Communauté Université Grenoble Alpes, 2018. [5] Faber, P. “Frames as a Framework for Terminology”. Handbook of Terminology. Ed. Hendrik J. Kockaert and Frieda Steurs. Amsterdam/Philadelphia: John Benjamins, 2015: 14-33.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,114 | 0,157 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,002 |
| Méta-épidémiologie (sens large) | 0,004 | 0,002 |
| Bibliométrie | 0,011 | 0,015 |
| Études des sciences et des technologies | 0,007 | 0,023 |
| Communication savante | 0,027 | 0,067 |
| Science ouverte | 0,012 | 0,023 |
| Intégrité de la recherche | 0,011 | 0,017 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,006 | 0,003 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».