Automated Clinical Practice Guideline Recommendations for Hereditary Cancer Risk Using Chatbots and Ontologies: System Description
Bibliographic record
Abstract
BACKGROUND: Identifying patients at risk of hereditary cancer based on their family health history is a highly nuanced task. Frequently, patients at risk are not referred for genetic counseling as providers lack the time and training to collect and assess their family health history. Consequently, patients at risk do not receive genetic counseling and testing that they need to determine the preventive steps they should take to mitigate their risk. OBJECTIVE: This study aims to automate clinical practice guideline recommendations for hereditary cancer risk based on patient family health history. METHODS: We combined chatbots, web application programming interfaces, clinical practice guidelines, and ontologies into a web service-oriented system that can automate family health history collection and assessment. We used Owlready2 and Protégé to develop a lightweight, patient-centric clinical practice guideline domain ontology using hereditary cancer criteria from the American College of Medical Genetics and Genomics and the National Cancer Comprehensive Network. RESULTS: The domain ontology has 758 classes, 20 object properties, 23 datatype properties, and 42 individuals and encompasses 44 cancers, 144 genes, and 113 clinical practice guideline criteria. So far, it has been used to assess >5000 family health history cases. We created 192 test cases to ensure concordance with clinical practice guidelines. The average test case completes in 4.5 (SD 1.9) seconds, the longest in 19.6 seconds, and the shortest in 2.9 seconds. CONCLUSIONS: Web service-enabled, chatbot-oriented family health history collection and ontology-driven clinical practice guideline criteria risk assessment is a simple and effective method for automating hereditary cancer risk screening.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".