پدافند الکترونیکی و سایبری

پدافند الکترونیکی و سایبری

ارائه ی یک مدل لب خوانی فارسی مبتنی بر ترکیب شبکه‌های پیچشی و ترنسفورمر

نوع مقاله : مقاله پژوهشی

نویسندگان
1 دانشجوی دکتری،دانشگاه یزد، یزد ، ایران
2 استاد،دانشگاه یزد، یزد ، ایران
3 دانشیار،دانشگاه یزد، یزد ، ایران
چکیده
گفتار رایج‌ترین شکل ارتباط بین انسان‌ها است و شامل درک اطلاعات صوتی و تصویری می‌شود. لب‌خوانی فرایندی است که برای تشخیص گفتار از روی حرکات لب گوینده استفاده می‌شود. ازآنجاکه محرمانگی گفتار در محیط‌های امنیتی به دلایل مسائل شنود بسیار اهمیت دارد؛ بنابراین لب‌خوانی به یکی از مهم‌ترین مسائل امروزه تبدیل شده است. در این مقاله روشی نوین مبتنی بر ترکیب شبکه‌های پیچشی و شبکه‌ی ترنسفورمر در تشخیص گفتار از تصویر حرکات لب افراد مختلف ارائه می‌شود. در این ساختار ویژگی‌ها با استفاده از شبکه‌های پیچشی استخراج می‌گردند. سپس این ویژگی‌ها به‌عنوان ورودی به ساختار ترنسفورمر پیشنهادی برای تولید متن عبارت گفتاری در خروجی منتقل می‌شوند. الگوریتم پیشنهادی بر روی مجموعه‌دادگان پاوید در زبان فارسی ارزیابی گردیده است. نوآوری روش پیشنهادی بازشناسی حرکات لب در سطح واحدهای واج یا ویزم است. این رویکرد امکان تعمیم‌پذیری بهتر و بازسازی گفتار پیوسته را فراهم می‌سازد. علاوه بر این آزمایش‌ها نشان می‌دهد رویکرد پیشنهادی زمان آموزش را کاهش می‌دهد و عملکرد سامانه را ارتقا می‌بخشد. دقت روش پیشنهادی در تشخیص واج‌های داده‌های آزمون، ۸۹ درصد است که نسبت به مدل‌های دیگر دارای دقت بالاتری است.
کلیدواژه‌ها
موضوعات

عنوان مقاله English

A Persian Lipreading Model Based on the Combination of CNN and Transformer Networks

نویسندگان English

zhina shahidi zandi 1
Vali Derhami 2
ali mohammad latif 2
mohammad taghi sadeghi 3
1 PhD Student, Yazd University, Yazd, Iran
2 Professor. Yazd University, Yazd, Iran
3 Associate Professor, Yazd University, Yazd, Iran
چکیده English

Speech is the primary medium of human communication and involves the interpretation of both auditory and visual information. Lipreading, or visual speech recognition, aims to infer spoken content by analyzing lip movements. With the growing importance of speech confidentiality in security-sensitive environments and the increasing vulnerability to acoustic eavesdropping, lipreading has attracted significant research attention in recent years. This paper presents a novel visual speech recognition framework that integrates convolutional neural networks (CNNs) with Transformer architectures to recognize speech from lip movement images of multiple speakers. In the proposed framework, CNNs extract discriminative spatial features, which are subsequently processed by a Transformer-based model to generate the corresponding textual output. The proposed method is evaluated on the Persian PAVID dataset. A key contribution of this work is viseme-level lip movement recognition, which improves generalization capability and enables continuous speech reconstruction. Experimental results demonstrate that the proposed approach reduces training time while enhancing overall recognition performance, achieving an accuracy of 89% in phoneme recognition on the test set and outperforming existing state-of-the-art methods.

کلیدواژه‌ها English

Lip reading
Visual speech recognition
convolutional neural networks (CNNs)
Feature extraction
Transformer

 

Smiley face

 

 

 

S. Jeon, A. Elsharkawy, and M. S. Kim, “Lipreading Architecture Based on Multiple Convolutional Neural Networks for Sentence-Level Visual Speech Recognition,” Sensors, vol. 22, no. 1, Art. no. 1, Jan. 2022. https://doi.org/10.3390/s22010072.
[2]           A. Bastanfard, M. Aghaahmadi, A. A. kelishami, M. Fazel, and M. Moghadam, “Persian Viseme Classification for Developing Visual Speech Training Application,” in Advances in Multimedia Information Processing - PCM 2009, P. Muneesawang, F. Wu, I. Kumazawa, A. Roeksabutr, M. Liao, and X. Tang, Eds., Berlin, Heidelberg: Springer, 2009, pp. 1080–1085. https://doi.org/10.1007/978-3-642-10467-1_104.
[3]           Y. Pei and H. Zha, “Stylized synthesis of facial speech motions,” Computer Animation and Virtual Worlds, vol. 18, no. 4–5, pp. 517–526, Sep. 2007.  https://doi.org/10.1002/cav.186.
[4]           T. Zhang, L. He, X. Li, and G. Feng, “Efficient End-to-End Sentence-Level Lipreading with Temporal Convolutional Networks,” Applied Sciences, vol. 11, no. 15, Art. no. 15, Jan. 2021. https://doi.org/ 10.3390/app11156975.
[5]           H. Mcgurk and J. Macdonald, “Hearing lips and seeing voices,” Nature, vol. 264, no. 5588, pp. 746–748, Dec. 1976. https://doi.org/10.1038/264746a0.
[6]           N. Akhter and A. Chakrabarty, “A Survey-based Study on Lip Segmentation Techniques for Lip Reading Applications,” 2016. Accessed: Jun. 22, 2025. [Online]. Available: https://www.semanticscholar.org/paper/A-Survey-based-Study-on-Lip-Segmentation-Techniques-Akhter-Chakrabarty/457a8efb4d9b71770ccddf0c74b88179a097e500
[7]           D. Yu, “The application of manifold based visual speech units for visual speech recognition,” PhD Thesis, Dublin City University, 2008. Accessed: Jun. 22, 2025. [Online]. Available: https://doras.dcu.ie/598/
[8]           W. C. Yau, D. K. Kumar, and S. P. Arjunan, “Visual recognition of speech consonants using facial movement features,” Integr. Comput.-Aided Eng., vol. 14, no. 1, pp. 49–61, Jan. 2007.
[9]           E. Petajan, B. Bischoff, D. Bodoff, and N. M. Brooke, “An improved automatic lipreading system to enhance speech recognition,” in Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, in CHI ’88. New York, NY, USA: Association for Computing Machinery, May 1988, pp. 19–25. https://doi.org/10.1145/57167.57170.
[10]         A. Maroosi, E. Zabbah, and H. Ataei Khabbaz, “Network Intrusion Detection using a combination of artificial neural networks in a hierarchical manner,” Electronic and Cyber Defense, vol. 8, no. 1, pp. 89–99, 2020. (In Persian)
[11]         A. Dolat Khah, M. Asadpour, R. Hashempour, and behnam dorostkar, “Violent behavior detection in surveillance cameras using convolutional and memory neural networks,” Electronic and Cyber Defense, vol. 12, no. 4, p., 2025. (In Persian)
[12]         S. Hourali, F. Hourali, and A. Pakzad, “Cyber Threat Information Extraction using Deep Learning and Knowledge Representation,” Electronic and Cyber Defense, vol. 13, no. 2, p., 2025. (In Persian)
[13]         Y. M. Assael, B. Shillingford, S. Whiteson, and N. de Freitas, “LipNet: End-to-End Sentence-level Lipreading,” Dec. 16, 2016, arXiv: arXiv:1611.01599. https://doi.org/ 10.48550/arXiv.1611.01599.
[14]         K. Xu, D. Li, N. Cassimatis, and X. Wang, “LCANet: End-to-End Lipreading with Cascaded Attention-CTC,” Mar. 13, 2018, arXiv: arXiv:1803.04988. https://doi.org/ 10.48550/arXiv.1803.04988.
[15]         M. A. Abrar, A. N. M. N. Islam, M. M. Hassan, M. T. Islam, C. Shahnaz, and S. A. Fattah, “Deep Lip Reading-A Deep Learning Based Lip-Reading Software for the Hearing Impaired,” in 2019 IEEE R10 Humanitarian Technology Conference (R10-HTC) (47129), Nov. 2019, pp. 40–44. https://doi.org/10.1109/R10-HTC47129.2019.9042439.
[16]         D. Parekh, A. Gupta, S. Chhatpar, A. Y. Kumar, and M. Kulkarni, “Lip Reading Using Convolutional Auto Encoders as Feature Extractor,” May 31, 2018, arXiv: arXiv:1805.12371. https://doi.org/10.48550/arXiv.1805.12371.
[17]         M. Riva, M. Wand, and J. Schmidhuber, “Motion Dynamics Improve Speaker-Independent Lipreading,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 2020, pp. 4407–4411. https://doi.org/10.1109/ICASSP40776.2020.9053535.
[18]         H. Huang, Ch. Song, J. Ting, T. Tian, Ch. Hong, Zh. Di, D. Gao, “A Novel Machine Lip Reading Model,” Procedia Computer Science, vol. 199, pp. 1432–1437, Jan. 2022. https://doi.org/10.1016/j.procs.2022.01.181.
[19]         H. Wang, G. Pu, and T. Chen, “A Lip Reading Method Based on 3D Convolutional Vision Transformer,” IEEE Access, vol. 10, pp. 77205–77212, 2022. https://doi.org/10.1109/ACCESS.2022.3193231.
[20]         A. Gholipour, H. Mohammadzade, A. Ghadami, and A. Taheri, “Automatic Lip Reading of Persian Words by a Robotic System Using Deep Learning Algorithms,” Iran J Sci Technol Trans Electr Eng, vol. 48, no. 4, pp. 1519–1538, Dec. 2024. https://doi.org/10.1007/s40998-024-00756-4.
[21]         M. Hedayatipour, Y. Shekofteh, and M. E. Moghaddam, “PAVID-CV s: Persian Audio-Visual Database of CV syllables,” in 2021 29th Iranian Conference on Electrical Engineering (ICEE), May 2021, pp. 470–473. https://doi.org/10.1109/ICEE52715.2021.9544268.
[22]         A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, “Attention Is All You Need,” Aug. 02, 2023, arXiv: arXiv:1706.03762. https://doi.org/10.48550/arXiv.1706.03762.
[23]         K. Paleček, “Extraction of Features for Lip-reading Using Autoencoders,” in Speech and Computer, A. Ronzhin, R. Potapova, and V. Delic, Eds., Cham: Springer International Publishing, 2014, pp. 209–216. https://doi.org/10.1007/978-3-319-11581-8_26.
[24]         D. E. King, davisking/dlib-models. (Oct. 29, 2025). C++. Accessed: Oct. 31, 2025. [Online]. Available: https://github.com/davisking/dlib-models
[25]         “Facial landmarks with dlib, OpenCV, and Python - PyImageSearch.” Accessed: Oct. 31, 2025. [Online]. Available: https://pyimagesearch.com/2017/04/03/facial-landmarks-dlib-opencv-python/

  • تاریخ دریافت 01 اسفند 1404
  • تاریخ بازنگری 22 اردیبهشت 1405
  • تاریخ پذیرش 16 خرداد 1405
  • تاریخ انتشار 03 تیر 1405