Exploring the Role of Terminology in SNOMED CT Definitions: Challenges and Solutions
Bibliographic record
Abstract
This work aims to show how insights from Terminology Science can foster interoperability and more effective knowledge representation, organisation, and sharing in the biomedical field. Definitions are one of the key forms of concept representation in Terminology and allow a given concept to be identified within a concept system while also enabling its differentiation from other related concepts [1]. They are also important to ensure conceptual and linguistic consistency of specialised terminology. Furthermore, definitions are an essential tool for those who need to access (or work on) a biomedical terminological resource but are not domain experts, such as translators (and patients). However, producing accurate and understandable definitions in natural language (i.e. text format) is a difficult and time-consuming task. Terminology can provide fundamental support in this process. SNOMED CT concepts are defined in three different ways [2]: a) the Fully Specified Name (FSN); b) a formal concept definition; and c) textual definitions. These three types of definitions can be useful for the work carried out by the interdisciplinary teams responsible for translating SNOMED CT content into the various languages of the SNOMED International Member Countries. However, professional translators without medical expertise will have greater difficulty understanding all SNOMED CT concepts starting from the FSN and/or their formal relationships. Moreover, the formal representation of concepts may contain errors which can only be detected by translators with sufficient domain knowledge. Therefore, natural language definitions and textual resources that provide information about the contextual use of a term are key additional tools. However, for the time being, textual definitions in SNOMED CT are optional and only provided for a limited number of concepts, when there is a requirement for additional detail. Having a language-independent, logically consistent, and interoperable conceptual foundation can help optimise the drafting process without jeopardising the added value supplied by linguistic diversity. Anchored in a double-dimensional approach to Terminology [3], which considers, within a given domain, not only the systematic and semi-automatic processing and analysis of specialised texts and their terms but also a (semi)formal representation of the core concepts and their respective organisation into systems, we propose a methodology for textual definition drafting based on the formal definitions in SNOMED CT. One of the examples analysed in the case study refers to |Single-port laparoscopic cholecystectomy (procedure)|, a primitive concept without a natural language definition proposal. To support the drafting of the latter, the formal definition was analysed, as well as its alignment with the existing FSN. However, it became clear that from a terminological standpoint, the concept |Single-port laparoscopic cholecystectomy (procedure)| currently has the same formal definition as its ‘parent’ concept |Laparoscopic cholecystectomy (procedure)| - a fully defined concept -, even though they are, in fact, two different concepts. Furthermore, the delimiting characteristic of the concept under analysis is not represented in the formal definition, despite its presence in the FSN (“single-port”). The proposal to incorporate the essential characteristic of |Single-port laparoscopic cholecystectomy (procedure)| into the formal definition, thereby enabling the subsequent textual definition to reflect that, is based on EndoTerm [4], which combines an Aristotelian-based, specific difference approach in the main hierarchies with a customised set of non-hierarchical concept relations, mostly based on SNOMED CT and UMLS, as well as with a categorial structure, to prevent logical errors. The concept systems within EndoTerm were encoded as a formal ontology, thereby enabling interoperability. An example is provided in Figure 1 for the micro-concept system of |Laparoendoscopic single-site total hysterectomy (procedure)|, a surgical procedure with the same surgical approach as |Single-port laparoscopic cholecystectomy (procedure)|. After the formal definition has been stabilised, a definitional template-like structure was created to optimise the textual definition drafting process and enhance consistency. The notion of definitional templates has been developed in terminology work within Frame-based Terminology [5]. In this work, this definitional template approach is taken further via the incorporation of ontologies and the use of formal SNOMED CT definitions. The definitional template in question can be seen in Figure 2 and based on it, a natural language definition proposal for |Laparoendoscopic single-site total hysterectomy (procedure)| would be “Procedure consisting of the excision of the uterus and cervix using a laparoscope to access the site via a minimally invasive surgical approach with a single skin incision”. Following this approach, only a minor, indispensable change would be required for |Single-port laparoscopic cholecystectomy (procedure)|, namely the body structure being excised. While ideally, the three concept definitions in SNOMED CT should be aligned, in order to facilitate the automation of the textual definition drafting process, this is not always the case, as the analysis of other examples has shown. Given the growing trend in incorporating textual definitions into biomedical terminological resources, as well as the importance that such natural language definitions have for translators and, potentially, for other user groups (namely patients), it is believed that Terminology – with its double-dimensional nature – and ontologies can bring added value from a methodological standpoint, providing consistency to both formal and linguistic aspects of the SNOMED CT definitions and contributing to a more successful alignment between them. References: [1] ISO. ISO 1087:2019. Terminology work and terminology science — Vocabulary, Geneva, 2019. [2] SNOMED International. SNOMED CT Editorial Guide, 2022. URL: https://confluence.ihtsdotools.org/display/DOCEG [3] Santos, C. and Costa, R. “Domain Specificity: Semasiological and Onomasiological Knowledge Representation”. Handbook of Terminology. Ed. Hendrik J. Kockaert and Frieda Steurs. Amsterdam / Philadelphia: John Benjamins, 2015: 153-179. [4] Carvalho, S. A terminological approach to knowledge organization within the scope of endometriosis: the EndoTerm project. PhD, Univ. NOVA de Lisboa - FCSH / Communauté Université Grenoble Alpes, 2018. [5] Faber, P. “Frames as a Framework for Terminology”. Handbook of Terminology. Ed. Hendrik J. Kockaert and Frieda Steurs. Amsterdam/Philadelphia: John Benjamins, 2015: 14-33.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.114 | 0.157 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.004 | 0.002 |
| Bibliometrics | 0.011 | 0.015 |
| Science and technology studies | 0.007 | 0.023 |
| Scholarly communication | 0.027 | 0.067 |
| Open science | 0.012 | 0.023 |
| Research integrity | 0.011 | 0.017 |
| Insufficient payload (model declined to judge) | 0.006 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".