PENGGUNAAN FITUR WARNA DAN TEKSTUR \nUNTUK CONTENT BASED IMAGE RETRIEVAL CITRA BUNGA
Bibliographic record
Abstract
Pencarian gambar berdasarkan gambar pada database, seringkali dilakukan untuk mengatasi duplikasi pada suatu karya. Content Based Image Retrieval (CBIR) Citra Bunga adalah engine pada komputer untuk melakukan pencarian gambar berdasarkan gambar pada database. Penelitian pada Content Based Image Retrieval (CBIR) Citra Bunga telah dilakukan oleh banyak peneliti. Permasalahan terjadi ketika memilih metode pendekatan seperti preprocessing, ekstraksi fitur dan similarity measure pada CBIR Citra Bunga. Pendekatan yang tidak sesuai dengan data yang diuji, tidak akan memberikan hasil yang optimal. Untuk mengetahui tingkat keberhasilan pendekatan yang digunakan pada CBIR Citra Bunga, digunakan perhitungan nilai precision. Pada penelitian ini, dataset yang akan digunakan adalah dataset Oxford Flower 17. Berdasarkan penelitian sebelumnya, untuk mendapatkan nilai precision yang lebih baik, penelitian ini akan menggunakan ekstraksi fitur warna Hue Saturation Value (HSV), ekstraksi fitur tekstur Gray Level Co-occurrence Matrix (GLCM), dan gabungan kedua fitur dengan pendekatan histogram. Pada penelitian CBIR Citra Bunga ini, terdapat tiga proses yaitu segmentasi menggunakan thresholding, proses ekstraksi fitur, dan pengukuran tingkat kemiripan citra dengan Euclidean Distance. Pengujian pada sistem dilakukan berdasarkan citra yang tersegmentasi dan tidak tersegmentasi. Pengujian sistem dengan hasil Mean Average Precision (MAP) terbesar dihasilkan oleh proses ekstraksi fitur GLCM tidak tersegmentasi sebesar 87,32%, dan untuk nilai MAP terbesar pada citra tersegmentasi dihasilkan pada proses ekstraksi fitur HSV sebesar 83,35%. \n \n \nKata kunci: Content Based Image Retrieval, ekstraksi fitur HSV, ekstraksi fitur GLCM, thresholding, Euclidean Distance, Mean Average Precision (MAP);--- \nSearching images based on images in the database, often done to overcome duplication of a work. Content Based Image Retrieval (CBIR) Flower Image is the engine on the computer To perform image-based image search on the database. Research on Content Based Image Retrieval (CBIR) Flower Image has been done by many researchers. Problems occur when choosing approaches such as preprocessing, feature extraction and similarity measure in CBIR Flower Image. Approaches which don't correspond with the data image test, would not provide optimal results. To know the success rate of approach used in CBIR Flower Image, the calculation of precision value is used. In this study, the dataset that will be used is dataset Oxford Flower 17. Based on previous research, to get better precision value, this research will use Hue Saturation Value (HSV) feature extraction, feature extraction of Gray Level Co-occurrence Matrix (GLCM) texture, and combination of both features with histogram approach. In this research, there are three processes: segmentation using thresholding, feature extraction process, and measurement of image similarity level with Euclidean Distance. For testing the system, is based on segmented image and non-segmented image. The result of the largest Mean Average Precision (MAP) produced in this study, resulted from the process of unsegmented image by the GLCM feature extraction of 87.32%, and for the largest MAP value in the segmented image produced by the HSV feature extraction process of 83.35%. \n \n \nKeywords: Content Based Image Retrieval, feature extraction HSV, feature extraction GLCM, thresholding, Euclidean Distance, Mean Average Precision (MAP)
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.028 | 0.024 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".