Facial Expression Recognition with ViT Considering All Tokens towards More Informative Self-attention Outputs
DOI:
https://doi.org/10.54097/hset.v41i.6745Keywords:
component; Facial Expression Recognition; ViT.Abstract
Currently few works have been done to apply Vision Transformer (ViT) on facial expression recognition (FER) successfully. In this paper, we put forward a novel idea of processing the outputs from the multi-head attention in ViT by passing through a global average pooling layer, and accordingly design 2 network architectures, namely ViTTL and ViTEH. These 2 models, especially the latter one, show more strength in recognition of local patterns, which normal ideas hold that it is a weakness of ViT. Extensive experiments are conducted on the benchmark dataset FER-2013 with popular state-of-the-art models employed as baselines for comparison. The results demonstrate the superiority of the 2 new models in the realm of FER.
Downloads
References
A. Haag, S. Goronzy, P. Schaich, and J. Williams, ‘Emotion Recognition Using Bio-sensors: First Steps towards an Automatic System’, in Affective Dialogue Systems, vol. 3068, E. André, L. Dybkjær, W. Minker, and P. Heisterkamp, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2004, pp. 36–48.
COMMUNICATION THEORY. LONDON: ROUTLEDGE, 2017. Accessed: Jul. 11, 2022.
A. Dosovitskiy et al., ‘An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale’. arXiv, Jun. 03, 2021.
Y. Tang, ‘Deep Learning using Linear Support Vector Machines’. arXiv, Feb. 21, 2015. Accessed: Jun. 26, 2022.
S. Zhao, ‘Feature Selection Mechanism in CNNs for Facial Expression Recognition’, p. 12.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, ‘ImageNet classification with deep convolutional neural networks’, Commun. ACM, vol. 60, no. 6, pp. 84–90, May 2017.
F. Xue, Q. Wang, and G. Guo, ‘TransFER: Learning Relation-aware Facial Expression Representations with Transformers’, in 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, Oct. 2021, pp. 3581–3590.
S. Kim, J. Nam, and B. C. Ko, ‘Facial Expression Recognition Based on Squeeze Vision Transformer’, Sensors, vol. 22, no. 10, p. 3729, May 2022, doi: 10.3390/s22103729.
J. L. Ba, J. R. Kiros, and G. E. Hinton, ‘Layer Normalization’. arXiv, Jul. 21, 2016.
M. Lin, Q. Chen, and S. Yan, ‘Network In Network’. arXiv, Mar. 04, 2014.
I. J. Goodfellow et al., ‘Challenges in Representation Learning: A Report on Three Machine Learning Contests’, in Neural Information Processing, vol. 8228, M. Lee, A. Hirose, Z.-G. Hou, and R. M. Kil, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2013, pp. 117–124.
I. Loshchilov and F. Hutter, ‘Decoupled Weight Decay Regularization’. arXiv, Jan. 04, 2019.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.







