Noise reduction and best band selection techniques for improving classification results using hyperspectral data: application to lithological mapping in Canada's Arctic
Bibliographic record
Abstract
AbstractTwo problems in using hyperspectral data are the effects of noise and redundancy of information (too many channels) on classification results. This paper introduces two methods for dealing with these problems using a hyperspectral dataset over southern Baffin Island in Canada's Arctic region. This paper shows how classification results using matched filtering (MF) are improved on the modified datasets for identifying various lithologies. The noise-reduced dataset produced using the inverse minimum noise fraction (MNF) transform provides the most accurate classifications, followed by the dataset comprising the best bands determined through eigenvector analysis. The supervised approach presented in this paper in which training areas are extracted using a combination of visual analysis of MNF component images, analysis of existing geology, and field observations followed by match filtering produce useful spectral maps to assist in focusing field mapping activities or as stand-alone maps that provide lithologic information even in the absence of field mapping.Un des problèmes rencontrés dans l'utilisation des données hyperspectrales provient des effets du bruit et de la redondance de l'information (trop de bandes) sur les résultats de classification. Le présent article présente deux méthodes pour pallier ces problèmes à l'aide d'un ensemble de données hyperspectrales acquises dans le sud de l'île de Baffin, dans l'Arctique canadien. Cet article montre comment les résultats de classification peuvent être améliorés à l'aide de la technique de filtrage adapté MF (« matched filtering ») dans le cas des ensembles de données modifiés pour l'identification des différentes lithologies. L'ensemble de données au bruit réduit utilisant l'inverse de la transformée MNF (« minimum noise fraction ») donne les classifications les plus précises suivies par les ensembles de données constitués des meilleures bandes déterminées par analyse des vecteurs propres. L'approche dirigée présentée dans cet article, dans laquelle les zones d'entraînement sont extraites en combinant l'analyse visuelle des images dérivées de la MNF, l'analyse de la géologie existante et des observations sur le terrain, suivie par une analyse MF, donne des cartes spectrales utiles pour orienter les activités de cartographie sur le terrain ou qui peuvent être utilisées tout simplement comme cartes apportant une information sur la lithologie, même en l'absence de cartes de terrain.[Traduit par la Rédaction]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".