Reducing Execution Time of Pixel-Based Machine Learning Classification Algorithms Using Parallel Processing Concept
Bibliographic record
Abstract
Parallel processing is essential in machine learning to meet the computational requirements resulting from the complexity of algorithms and the size of the dataset, by taking advantage of the computational resources of parallel processing that can distribute computational operations across multiple processors. Which contributes to significant improvements in performance and time efficiency. This research demonstrated the impact of parallel processing on the performance and time efficiency of machine learning for pixel-based image classification techniques. The methodology includes pre-processing the Oxford IIIT Pet dataset, from which 4 cat images were selected. The performance of two supervised machine learning classifiers, decision tree, and random forest (10, 100, 500, and 1000 trees) were compared and implemented in two ways with and without parallel processing. The data is split in two ways: the first is by splitting the data by 70% for training data and 30% for testing data and the second is by cross-validation by splitting the data into four folds. The research aims to compare the accuracy and timely scales of machine learning models with and without parallel processing. The results showed a strong predictive power of the algorithms with an accuracy of 97.5%, while the training times were significantly reduced in parallel from 88.83 to 15.88 seconds for the RF100 model for image no. 2. This reflects the effectiveness of parallel processing in improving the performance of machine-learning models for pixel-based image classification. The proposed system was programmed using MATLAB 2021 language tools.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".