Analysis of Lecxico-Semantic Relations of Punjabi Shahmukhi Nouns: A Corpus Based Study
Bibliographic record
Abstract
The current study is an effort in the development of Lecxico-semantic relations among Punjabi Shahmukhi nouns. Semantic relations are those nets, which are found among nouns on the bases of word meanings. Development of semantic nets is taken as a key part while developing WordNet of any language. The WordNet of Punjabi Shahmukhi is not developed yet. The digital exposure and progress of Punjabi Shahmukhi is very slow in comparison to other languages of the world. The present study explores the kind of semantic relations found among the nouns of Punjabi Shahmukhi. WordNet organizes words on the basis of word meanings rather than word forms. WordNet of English includes four open class categories including; nouns, verbs, adverbs and adjectives, but present study is limited to the analysis of nouns. A corpus of 2 million words of Punjabi Shahmukhi was taken from different sources. Then, it was POS tagged and a list of 846 nouns was generated. Then, each noun was analyzed individually to develop its Lecxico-semantic relations including: synonymy, antonymy, meronyms, holonymy, hyponymy, hypernymy, singular, plural, masculine, feminine and HAS a part. The present research is significant and useful in the development of WordNet for Punjabi Shahmukhi. With the development of WordNet, it will be possible to run digital applications in Punjabi Shahmukhi including: machine translation, information retrieval, querying archive and report generation to automatic speech recognition, data mining, read aloud, robotics and many more. On the other hand, WordNet will help to maintain an international status for Punjabi Shahmukhi.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.026 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".