CLustre: semi-automated lineament clustering for palaeo-glacial reconstruction
Bibliographic record
Abstract
Datasets containing large numbers (>10,000) of glacial lineaments are increasingly being mapped from remotely sensed data in order to develop a palaeo-glacial reconstruction or ”inversion”. The palimpsest landscape presents a complex record of past ice flow and deconstructing this information into a logical history is an involved task. One stage in this process requires the identification of sets of genetically linked lineaments that can form the basis of a reconstruction. \n \nThis paper presents a semi-automated algorithm, CLustre, for lineament clustering that uses a locally adaptive, region growing, methodology. After outlining the algorithm, it is tested on synthetic datasets that simulate parallel and orthogonal cross-cutting lineaments, encompassing 1,500 separate classifications. Results show robust classification in most scenarios, although parallel overlap of lineaments can cause false positive classification unless there are differences in lineament length. Case studies for Dubawnt Lake and Victoria Island, Canada, are presented and compared to existing datasets. For Dubawnt Lake 9 out of 14 classifications directly match incorporating 89% of lineaments. For Victoria Island 57 out of 58 classifications directly match incorporating 95% of lineaments. Differences are related to small numbers of unclassified lineaments and parallel cross-cutting lineaments that are of a similar length. \nCLustre enables the automated, repeatable, assignment of lineaments to flow sets using defined user criteria. This is important as qualitative visual interpretation may introduce bias, potentially weakening the testability of palaeo-glacial reconstructions. In addition, once classified, summary statistics of lineament clusters can be calculated and subsequently used during the reconstruction process.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".