Quantifying Aflatoxin B1 Contamination Levels in Almonds Using Hyperspectral Imaging Utilizing Gaussian Process and Support Vector Regression
Bibliographic record
Abstract
Almonds are prone to infection by Aspergillus fungi in warm and humid environments, which produce aflatoxin B1 (AFB1) in secondary metabolism. AFB1 is a toxic substance and continuous consumption of AFB1-contaminated almonds causes serious health problems. The current AFB1 measure in almonds is destructive, labor-intensive, costly, and inapplicable for industrial inline application. This research study employed hyperspectral images within the wavelength range of 900 to 1700 nm to explore the potential of quantifying aflatoxin B1 contamination levels in almond kernels. In this experiment, a full spectra Gaussian Process Regression (GPR) model and a Support Vector Regression (SVR) model were developed to quantify artificially aflatoxin B1 contaminated almonds. Genetic Algorithm (GA) was used to select significant feature spectra from the hyperspectral image dataset. Then, the significant features spectra were used to develop a multispectral GPR and SVR model to quantify aflatoxin B1 in almonds. The GPR model achieved a coefficient of regression (R2) value of 0.966 for training, 0.934 for testing, and 0.93 for cross-validation using raw spectra. Also, the SVR model achieved R2 value of 0.954 for training. The study shows that pairing hyperspectral imaging with machine learning can accurately measure AFB1 in individual almond kernels. This could be valuable to implementing AFB1 quantification in quality control for industrial purposes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".