Signature Verification by Multi-Size Assembled-Attention with the Backbone of Swin-Transformer
Bibliographic record
Abstract
<title>Abstract</title> Handwritten signature verification is an indispensable means of identification in biometric information recognition and has broad application prospects and research significance in financial, judicial, and educational systems. With the advancement of signature forgery technology, people have higher requirements for the accuracy and efficiency of signature verification, and sophisticated convolutional networks capable of automatic feature extraction are gradually being applied to the field of handwriting recognition. However, these convolution methods still have the potential for improvement in recognition capability, generalization capability, and accuracy rate. This paper proposes a novel network model, Multi-Size Assembled-Attention Swin-Transformer network, to perform signature handwriting authenticity identification. The inputs to the network are signature images that are resized to multiple sizes, including (224, 224), (112, 112), (56, 56). Then, features within the same image are extracted using the self-attention mechanism in Swin-Transformer, and features between different images are also extracted with the cross-attention mechanism in Assembled-Attention Block, enabling signature feature information to interact within the same image and between different images. Also, Regularized Dropout strategy and adversarial method are implemented in the training stage. Therefore, our method considerably prompts the identification ability of the signature handwriting and obtains state-of-the-art performance, especially 57.1% and 50.4% improvement, in the situation of training in CEDAR and evaluation in Bengali and Hindi. Meantime, we evaluated the impact of the input images passing through the model's times on performance and found that the network achieves the optimal performance at times of four.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".