A green analytical method for fish species authentication based on Raman spectroscopy
Bibliographic record
Abstract
Fish mislabeling is a rampant global issue, damaging consumers economic benefits and trust in the fish industry and government authorities, as well as diminishing the efficacy of the sustainability measurement and management of fisheries. Although DNA barcoding as a gold standard method provides accurate identification of biological species of fish, this method is complicated and slow, and requires reagents and solvents. To develop a more rapid, easy-to-use, and environmentally-friendly method for fish species identification, we integrated the non-destructive Raman spectroscopy with chemometrics/machine learning for rapid and simple fish species authentication. Two Raman spectrometers (i.e., a portable Raman spectrometer and a benchtop confocal Raman spectrometer) were used and compared for their performance to identify 11 species of fish (i.e., 4 species of Salmonidae and 7 species of non-Salmonidae). Supervised chemometric/machine learning classification models were constructed based on a hierarchical classification principle to solve this 11-class identification problem. Both Raman spectrometers were able to differentiate Salmonidae from non-Salmonidae fish with close to 100% accuracy (i.e., first-hierarchical level). To further identify the fish to species level, the portable Raman spectrometer provided better accuracy (i.e., 93% and 93% accuracy for the Salmonidae group and non-Salmonidae group of fish identification, respectively) compared to the benchtop Raman spectrometer (i.e., 90% and 84% accuracy for the Salmonidae group and non-Salmonidae group of fish identification, respectively). The overall analytical time from sample to results can be completed within 5 min, much faster compared to the gold standard method. Moreover, the classification power of this Raman spectroscopy-based technique is expected to be improved with an increased spectral number of fish species and biological replicates in the Raman spectral library, as well with advanced machine learning algorithms. This rapid and reliable fish authentication method based on Raman spectroscopy will provide government laboratories and the fish industry another useful tool to routinely and frequently monitor the fish authenticity, and thus to protect consumers’ benefits and guarantee the efficacy of the fishery sustainability measurement and management strategies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".