Revisiting Inductive Bias in Vision Transformers: A Comparative Analysis of Transfer Learning and Training from Scratch on CIFAR-10

Authors

  • Arya Putatunda ap9857@srmist.edu.in Author
  • Pranshu Kumar Department of Computer Science and Engineering SRM Institute of Science and Technology, Ghaziabad, India Author
  • Praveen Department of Computer Science and Engineering SRM Institute of Science and Technology, Ghaziabad, India Author
  • Akash Karan Department of Computer Science and Engineering SRM Institute of Science and Technology, Ghaziabad, India Author

DOI:

https://doi.org/10.63503/acset.115

Keywords:

Vision Transformer, Inductive Bias, Transfer Learning, CIFAR-10, Deep Learning, Hybrid Architectures

Abstract

This paper is a comparative study of transfer learning and training from scratch on CIFAR-10, and of the relationship between inductive bias and vision transformers. Vision Transformers (ViT) are a relatively new form of potent alternative to Convolutional Neural Networks (CNNs) in the field of computer vision. However, unlike CNNs, ViTs lack strong inductive biases, such as locality or translation invariance and thus owe a lot to large-scale data. The paper will conduct a parallel comparison between training a Vision Transformer on the CIFAR-10 dataset in a pure vanilla environment and using a pre-trained model via transfer learning. A Lightweight hybrid vision transformer is an ad-hoc training and it has an accuracy of 72.82%. In comparison, when a ViT-B/16 is pre-trained on ImageNet and subsequently fine-tuned on CIFAR-10, the accuracy reaches 94.08% after only five epochs. The results indicate that the latter is not as efficient at using data as Vision Transformers trained directly; although trained on a large scale, it has a high representation learning capacity.

 

References

[1] Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998.

[2] A. Vaswani et al., “Attention is all you need,” in Advances in Neural Information Processing Systems, 2017, pp. 5998–6008.

[3] A. Dosovitskiy et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations (ICLR), 2021.

[4] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778.

[5] A. Krizhevsky, “Learning multiple layers of features from tiny images,” Univ. of Toronto, Tech. Rep., 2009.

[6] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convo-lutional neural networks,” in NIPS, 2012.

[7] H. Touvron et al., “Training data-efficient image transformers & distillation through at-tention,” in ICML, 2021.

[8] Z. Liu et al., “Swin Transformer: Hierarchical vision transformer using shifted windows,” in ICCV, 2021.

[9] C.-F. Chen et al., “CrossViT: Cross-attention multi-scale vision transformer for image classification,” in ICCV, 2021.

[10] J.-B. Cordonnier, A. Loukas, and M. Jaggi, “On the relationship between self-attention and convolutional layers,” in ICLR, 2020.

[11] M. Caron et al., “Emerging properties in self-supervised vision transformers,” in ICCV, 2021.

[12] K. He et al., “Masked autoencoders are scalable vision learners,” in CVPR, 2022.

[13] H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” in ICLR, 2018.

[14] S. Yun et al., “CutMix: Regularization strategy to train strong classifiers with localizable features,” in ICCV, 2019.

[15] I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in ICLR, 2019.

[16] D. Hendrycks and T. Dietterich, “Benchmarking neural network robustness to common corruptions and perturbations,” in ICLR, 2019.

[17] S. Kaur et al., “On the loss landscape of vision transformers,” arXiv preprint arXiv:2305.16723, 2023.

[18] M. Raghu et al., “Do vision transformers see like convolutional neural networks?,” in NeurIPS, 2021.

[19] A. Steiner et al., “How to train your ViT? Data, augmentation, and regularization in vision transformers,” arXiv preprint arXiv:2106.10270, 2021.

[20] S. Khan et al., “Transformers in vision: A survey,” ACM Computing Surveys, vol. 54, no. 10s, pp. 1–41, 2022.

[21] Y. Zhang et al., “A survey on hybrid vision transformers,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

[22] M. Arslanoglu and M. Ozturk, “Understanding texture bias in vision transformers,” Pattern Recognition Letters, vol. 167, pp. 10–17, 2023.

[23] J. Deng et al., “ImageNet: A large-scale hierarchical image database,” in CVPR, 2009.

Downloads

Published

2026-09-15

Conference Proceedings Volume

Section

Articles

How to Cite

Arya Putatunda, Pranshu Kumar, Praveen, & Akash Karan. (2026). Revisiting Inductive Bias in Vision Transformers: A Comparative Analysis of Transfer Learning and Training from Scratch on CIFAR-10 . Adroid Conference Series: Engineering and Technology, 2(3), 105-117. https://doi.org/10.63503/acset.115