Comparative Study of Data Efficiency in Vision Transformer and ResNet-18 Architectures: Using CIFAR-10 and TinyImageNet
Bibliographic record
Abstract
Deep learning algorithms for computer vision have been primarily based on architectures utilizing convolutional layers for feature extraction until 2020, when Dosovitskiy et al. proved that the Vision Transformer, an attention-based neural network outperforms many state-of-the-art convolutional networks of that time in several computer vision tasks. The architecture of vision transformers differ fundamentally from convolutional networks. Convolutional layers in convolutional networks excel at capturing local features in images, whereas vision transformers are better suited for learning global features that convolutional networks often miss. This advantage comes at the cost of data efficiency, which has limited the adoption of vision transformers until recently. This thesis compares the learning efficiency of two neural networks of similar complexity: ResNet-18 by He et al. and the Vision Transformer by Dosovitskiy et al. Both models are trained on varying fractions of the CIFAR-10 and TinyImageNet datasets. Canadian Institute for Advanced Research-10 (CIFAR-10) consists of 60,000 RGB images (32x32 pixels, 10 classes), while TinyImageNet contains 110,000 RGB images (64x64 pixels, 200 classes). Results are compared across epochs and different training dataset fractions. The results show that as the number of epochs increases, both architectures learn similarly, with ResNet-18 models performing slightly better. The observed differences likely stem from the size of the datasets used in the experiment, which is not enough for the Vision Transformer to outperform ResNet-18.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".