Artificial intelligence for target symptoms of Thai herbal medicine by web scraping
Bibliographic record
Abstract
Machine learning (ML) is implementing artificial intelligence (AI) research within medicine that has made dramatic progress in recent years. In addition to standard treatments, the role of complementary and alternative medicine should be mentioned. Traditional Thai medicine has received growing acceptance as a complementary approach to modern medicine by using local herbs. A vast amount of Thai herbal knowledge and information is freely available on the Internet. The reader must evaluate each website and decide to use trustworthy and appropriate information. This study aimed to acquire Thai herbal knowledge recorded in the Thai language system on the Internet by scraping websites using programming techniques. The knowledge was extracted with programming, and the types of Thai herbs were classified corresponding to target symptoms by the machine learning algorithm. The ML method organized the process when sufficient achievement was reached in order to give reliable and high accuracy results from the training data set. The validation of extracted knowledge was achieved by using the part-of-speech tag patterns analysis. This study showed that the programming and machine learning system was appropriate for obtaining and classifying Thai herbal medicines knowledge.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.005 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".