A Decision Support Tool to Analyze Food Properties from Near Infrared Spectroscopy*
Bibliographic record
Abstract
Near infrared spectroscopy (NIRS) is an analytical technique that is gaining popularity in the food industry due to its low operating costs, rapid analysis and non-destructive sample technique. Numerous studies have shown the relevance of NIR spectra analysis to determine certain quality attributes of food. This makes it attractive for use in quality control and continuous monitoring of food processing. However, the calibration process of NIR is difficult and time-consuming. Depending on the configuration of the NIR instrument, the sample to be analyzed and the attribute to be predicted, the analysis methods and techniques vary. This makes calibration a challenge for many manufacturers. This article aims to develop a decision support tool to assess food properties based on the analysis of selected features of NIR spectra. It intends to provide support to calibrate a predictive model based on NIR spectra. The methodology-based decision support tool was evaluated on cocoa bean samples. The tool suggested using the SG filter technique and PLSR machine learning model to predict the moisture and fat content of cocoa beans. The PLSR model with 4 components trained from 63 wavelengths obtained excellent results for the prediction of moisture content with an R2CV of 0.9 and an RMSEP of 0.28. While the PLSR model with 8 components trained from 166 wavelengths obtained satisfactory results for the prediction of fat content with an R<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup>CV of 0.93 and an RMSEP of 1.52.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.021 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".