DBARCT: Road Extraction Based on Double-Branch Architecture and Random Block Coding Transformer
Bibliographic record
Abstract
Although transformer models are main network architectures for the delineation of roads from remote sensing imagery, they have critical limitations due to their regular patch mechanism and inefficiency in local information learning. To address these limitations for enhanced road extraction, this letter presents a novel double-branch architecture and random block coding transformer (DBARCT), with the following contributions. First, to improve local spatial details’ learning, we integrate transformer with convolutional neural network (CNN) into a novel dual-branch encoder-decoder architecture, such that the resulting model is efficient at learning both the local edge information and the global context information that are highly complementary for accurate road extraction. Second, to additionally augment the learning of global contextual information, we integrate the regular patching approach in traditional transformer models with a new irregular patching approach, such that it can better capture the global spatial information correlations that might be ignored by the regular patching approach. Third, an array of tests was carried out to meticulously scrutinize the efficacy of the fundamental elements of the suggested model. The empirical findings reveal that the intersection over union (IoU) metric attained by the proposed methodology on the LRSNY dataset stands at 88.53%, thereby corroborating the efficacy and preeminence of our approach in tasks related to road extraction.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".