Object Surface Defect Segmentation using Vision Transformers and Hybrid Models

Authors

  • Santhya C Department of Computational Intelligence, SRM Institute of Science and Technology, SRM Nagar, Chennai, 603203, Tamil Nadu, India Author
  • VS Sneha Chowdary Department of Computational Intelligence, SRM Institute of Science and Technology, SRM Nagar, Chennai, 603203, Tamil Nadu, India Author
  • Beaulah Jeyavathana Department of Computational Intelligence, SRM Institute of Science and Technology, SRM Nagar, Chennai, 603203, Tamil Nadu, India Author

DOI:

https://doi.org/10.63503/acset.109

Keywords:

Surface Defect Segmentation, Vision Transformers, Hybrid CNN-Transformer Models, Industrial Quality Inspection, Deep Learning, Dice Coefficient, Mean Intersection over Union (mIoU)

Abstract

Detection of surface defects in the manufacturing industry is fundamental to maintain the quality of products and reduce waste production. Inspection processes are always affected by human error due to the nature of manual operations in manufacturing environments, particularly in cases of high production rates, while traditional machine vision systems cannot generalize over the complex structure of surface materials. In order to achieve the goal of automated surface defect detection, convolutional neural networks have been proposed. However, the local feature extraction abilities restrict these network from detecting long-distance dependencies in surface material structures and hence limits its ability to detect defect boundaries accurately. This research paper offers an empirical comparison between three deep learning frameworks for pixel-based surface defect detection: a baseline CNN architecture known as U-Net, Vision Transformer (ViT) based on patch-wise attention mechanism, and TransUNet that combines CNN and transformer architecture. The three deep learning models were compared following a unified approach of training and evaluation over three benchmark industrial surface defect segmentation datasets, i.e., MVTec AD, NEU Surface Defect Database, and DAGM 2007, where the Dice Coefficient and mean Intersection over Union are used as evaluation criteria. It is observed that U-Net performs best in terms of maximum mIoU of 0.75 at Epoch 26, ViT attains 0.72 with better generalization behavior at Epoch 41, and TransUNet achieves 0.67 at epoch 19. 

References

[1] A. Dosovitskiy et al., “An image is worth 16×16 words: Transformers for image recognition at scale,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2021.

[2] Z. Liu et al., “Swin Transformer: Hierarchical vision transformer using shifted windows,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021, pp. 10012–10022.

[3] O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in Proc. Med. Image Comput. Comput. Assist. Interv. (MICCAI), 2015, pp. 234–241.

[4] P. Bergmann et al., “The MVTec anomaly detection dataset,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2019, pp. 9592–9600.

[5] K. He et al., “Mask R-CNN,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2017, pp. 2961–2969.

[6] W. Wang et al., “Pyramid vision transformer: A versatile backbone for dense prediction,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021, pp. 568–578.

[7] Y. Zhang, R. Jin, and X. Xu, “Deep learning-based steel surface defect detection,” IEEE Access, vol. 9, pp. 123456–123468, 2021.

[8] S. Mei, H. Yang, and Z. Yin, “Anomaly detection using transformer-based feature reconstruction,” Pattern Recognit., vol. 128, p. 108676, 2022.

[9] J. M. Valanarasu et al., “Medical transformer: Gated axial-attention for medical image segmentation,” in Proc. MICCAI, 2021, pp. 36–46.

[10] X. Xu, Y. Ding, and J. Li, “Surface defect detection based on deep learning,” IEEE Trans. Ind. Inform., vol. 16, no. 8, pp. 5555–5563, 2020.

[11] Z. Zhou et al., “UNet++: A nested U-Net architecture for medical image segmentation,” IEEE Trans. Med. Imaging, vol. 39, no. 6, pp. 1856–1867, 2019.

[12] J. Chen et al., “TransUNet: Transformers make strong encoders for medical image segmentation,” arXiv preprint arXiv:2102.04306, 2021.

[13] J. Tao, Y. Sun, and Y. Zhang, “Surface defect detection method based on deep learning,” Sensors, vol. 18, no. 9, p. 2865, 2018.

[14] E. Xie et al., “SegFormer: Simple and efficient design for semantic segmentation with transformers,” in Advances in Neural Information Processing Systems (NeurIPS), 2021.

[15] Y. Li, L. Zhao, and X. Wang, “Attention-based deep learning for surface defect detection,” IEEE Access, vol. 8, pp. 12345–12356, 2020.

[16] Z. Zhou et al., “The DAGM 2007 dataset for defect detection,” in Proc. Int. Conf. Pattern Recognit. (ICPR), 2007.

[17] Y. Tian et al., “Deep learning for industrial inspection: A survey,” J. Manuf. Syst., vol. 56, pp. 453–471, 2020.

[18] J. Wang, H. Chen, and Y. Li, “Hybrid CNN–Transformer network for image segmentation,” IEEE Access, vol. 10, pp. 98765–98776, 2022.

[19] H. Cao et al., “Swin-UNet: U-Net-like pure transformer for medical image segmentation,” in Proc. Eur. Conf. Comput. Vis. Workshops (ECCV Workshops), 2022.

[20] M. Khan, S. Ahmed, and R. Ali, “Vision transformer-based surface defect detection for industrial inspection,” Appl. Sci., vol. 13, no. 4, p. 2100, 2023

Downloads

Published

2026-09-15

Conference Proceedings Volume

Section

Articles

How to Cite

Santhya C, VS Sneha Chowdary, & Beaulah Jeyavathana. (2026). Object Surface Defect Segmentation using Vision Transformers and Hybrid Models . Adroid Conference Series: Engineering and Technology, 2(3), 48-60. https://doi.org/10.63503/acset.109