CLustre: semi-automated lineament clustering for palaeo-glacial reconstruction
Bibliographic record
Abstract
Datasets containing large numbers (>10,000) of glacial lineaments are increasingly being mapped from remotely sensed data in order to develop a palaeo-glacial reconstruction or ”inversion”. The palimpsest landscape presents a complex record of past ice flow and deconstructing this information into a logical history is an involved task. One stage in this process requires the identification of sets of genetically linked lineaments that can form the basis of a reconstruction. \n \nThis paper presents a semi-automated algorithm, CLustre, for lineament clustering that uses a locally adaptive, region growing, methodology. After outlining the algorithm, it is tested on synthetic datasets that simulate parallel and orthogonal cross-cutting lineaments, encompassing 1,500 separate classifications. Results show robust classification in most scenarios, although parallel overlap of lineaments can cause false positive classification unless there are differences in lineament length. Case studies for Dubawnt Lake and Victoria Island, Canada, are presented and compared to existing datasets. For Dubawnt Lake 9 out of 14 classifications directly match incorporating 89% of lineaments. For Victoria Island 57 out of 58 classifications directly match incorporating 95% of lineaments. Differences are related to small numbers of unclassified lineaments and parallel cross-cutting lineaments that are of a similar length. \nCLustre enables the automated, repeatable, assignment of lineaments to flow sets using defined user criteria. This is important as qualitative visual interpretation may introduce bias, potentially weakening the testability of palaeo-glacial reconstructions. In addition, once classified, summary statistics of lineament clusters can be calculated and subsequently used during the reconstruction process.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.004 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.006 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".