Electronic and Cyber Defense

Electronic and Cyber Defense

A Persian Lipreading Model Based on the Combination of CNN and Transformer Networks

Document Type : Original Article

Authors
1 PhD Student, Yazd University, Yazd, Iran
2 Professor. Yazd University, Yazd, Iran
3 Professor, Yazd University, Yazd, Iran
4 Associate Professor, Yazd University, Yazd, Iran
Abstract
Speech is the primary medium of human communication and involves the interpretation of both auditory and visual information. Lipreading, or visual speech recognition, aims to infer spoken content by analyzing lip movements. With the growing importance of speech confidentiality in security-sensitive environments and the increasing vulnerability to acoustic eavesdropping, lipreading has attracted significant research attention in recent years. This paper presents a novel visual speech recognition framework that integrates convolutional neural networks (CNNs) with Transformer architectures to recognize speech from lip movement images of multiple speakers. In the proposed framework, CNNs extract discriminative spatial features, which are subsequently processed by a Transformer-based model to generate the corresponding textual output. The proposed method is evaluated on the Persian PAVID dataset. A key contribution of this work is viseme-level lip movement recognition, which improves generalization capability and enables continuous speech reconstruction. Experimental results demonstrate that the proposed approach reduces training time while enhancing overall recognition performance, achieving an accuracy of 89% in phoneme recognition on the test set and outperforming existing state-of-the-art methods.
Keywords
Subjects

  • Receive Date 09 February 2026
  • Revise Date 26 March 2026
  • Accept Date 20 April 2026
  • Publish Date 22 May 2026