MétaCan
Menu
← Back to cohort
Record W4393424768 · doi:10.5281/zenodo.7916716

Text2KGBench: A Benchmark for Ontology-Driven Knowledge Graph Generation from Text

2023· dataset· en· W4393424768 on OpenAlexaboutno aff
Nandana Mihindukulasooriya, Sanju Tiwari, Carlos F. Enguix, Kusum Lata

Bibliographic record

VenueZenodo (CERN European Organization for Nuclear Research) · 2023
Typedataset
Languageen
FieldComputer Science
TopicSemantic Web and Ontologies
Canadian institutionsnot available
Fundersnot available
KeywordsBenchmark (surveying)OntologyComputer scienceKnowledge graphGraphNatural language processingInformation retrievalArtificial intelligenceWorld Wide WebTheoretical computer sciencePhilosophyGeographyCartographyEpistemology

Abstract

fetched live from OpenAlex

This is the repository for ISWC 2023 Resource Track submission for Text2KGBench: Benchmark for Ontology-Driven Knowledge Graph Generation from Text. Text2KGBench is a benchmark to evaluate the capabilities of language models to generate KGs from natural language text guided by an ontology. Given an input ontology and a set of sentences, the task is to extract facts from the text while complying with the given ontology (concepts, relations, domain/range constraints) and being faithful to the input sentences. It contains two datasets (i) Wikidata-TekGen with 10 ontologies and 13,474 sentences and (ii) DBpedia-WebNLG with 19 ontologies and 4,860 sentences. An example An example test sentence: Test Sentence: {"id": "ont_music_test_n", "sent": "\"The Loco-Motion\" is a 1962 pop song written by American songwriters Gerry Goffin and Carole King."} An example of ontology: Ontology: Music Ontology Expected Output: { "id": "ont_k_music_test_n", "sent": "\"The Loco-Motion\" is a 1962 pop song written by American songwriters Gerry Goffin and Carole King.", "triples": [ { "sub": "The Loco-Motion", "rel": "publication date", "obj": "01 January 1962" },{ "sub": "The Loco-Motion", "rel": "lyrics by", "obj": "Gerry Goffin" },{ "sub": "The Loco-Motion", "rel": "lyrics by", "obj": "Carole King" },] } The data is released under a Creative Commons Attribution-ShareAlike 4.0 International (CC BY 4.0) License. The structure of the repo is as the following. Text2KGBench src: the source code used for generation and evaluation, and baseline benchmark the code used to generate the benchmark evaluation evaluation scripts for calculating the results baseline code for generating the baselines including prompts, sentence similarities, and LLM client. data: the benchmark datasets and baseline data. There are two datasets: wikidata_tekgen and dbpedia_webnlg. wikidata_tekgen Wikidata-TekGen Dataset ontologies 10 ontologies used by this dataset train training data test test data manually_verified_sentences ids of a subset of test cases manually validated unseen_sentences new sentences that are added by the authors which are not part of Wikipedia test unseen test unseen test sentences ground_truth ground truth for unseen test sentences. ground_truth ground truth for the test data baselines data related to running the baselines. test_train_sent_similarity for each test case, 5 most similar train sentences generated using SBERT T5-XXL model. prompts prompts corresponding to each test file unseen prompts unseen prompts for the unseen test cases Alpaca-LoRA-13B data related to the Alpaca-LoRA model llm_responses raw LLM responses and extracted triples eval_metrics ontology-level and aggregated evaluation results unseen results results for the unseen test cases llm_responses raw LLM responses and extracted triples eval_metrics ontology-level and aggregated evaluation results Vicuna-13B data related to the Vicuna-13B model llm_responses raw LLM responses and extracted triples eval_metrics ontology-level and aggregated evaluation results dbpedia_webnlg DBpedia Dataset ontologies 19 ontologies used by this dataset train training data test test data ground_truth ground truth for the test data baselines data related to running the baselines. test_train_sent_similarity for each test case, 5 most similar train sentences generated using SBERT T5-XXL model. prompts prompts corresponding to each test file Alpaca-LoRA-13B data related to the Alpaca-LoRA model llm_responses raw LLM responses and extracted triples eval_metrics ontology-level and aggregated evaluation results Vicuna-13B data related to the Vicuna-13B model llm_responses raw LLM responses and extracted triples eval_metrics ontology-level and aggregated evaluation results This benchmark contains data derived from the TekGen corpus (part of the KELM corpus) [1] released under CC BY-SA 2.0 license and WebNLG 3.0 corpus [2] released under CC BY-NC-SA 4.0 license. [1] Oshin Agarwal, Heming Ge, Siamak Shakeri, and Rami Al-Rfou. 2021. Knowledge Graph Based Synthetic Corpus Generation for Knowledge-Enhanced Language Model Pre-training. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3554–3565, Online. Association for Computational Linguistics. [2] Claire Gardent, Anastasia Shimorina, Shashi Narayan, and Laura Perez-Beltrachini. 2017. Creating Training Corpora for NLG Micro-Planners. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 179–188, Vancouver, Canada. Association for Computational Linguistics.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.034
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.044
Threshold uncertainty score0.146

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0040.034
Meta-epidemiology (narrow)0.0040.001
Meta-epidemiology (broad)0.0010.003
Bibliometrics0.0070.006
Science and technology studies0.0010.001
Scholarly communication0.0030.005
Open science0.0050.004
Research integrity0.0030.002
Insufficient payload (model declined to judge)0.0440.023

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.074
GPT teacher head0.283
Teacher spread0.210 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueZenodo (CERN European Organization for Nuclear Research)→Same topicSemantic Web and Ontologies→French-language works237,207→