Lu Yuanyao, Li Kexin
School of Information, North China University of Technology, Beijing 100144, China.
Math Biosci Eng. 2022 Sep 15;19(12):13526-13540. doi: 10.3934/mbe.2022631.
With the development of deep learning and artificial intelligence, the application of lip recognition is in high demand in computer vision and human-machine interaction. Especially, utilizing automatic lip recognition technology to improve performance during social interactions for those hard of hearing, and pronunciation is one of the most promising applications of artificial intelligence in medical healthcare and rehabilitation. Lip recognition means to recognize the content expressed by the speaker by analyzing dynamic motions. Presently, lip recognition research mainly focuses on the algorithms and computational performance, but there are relatively few research articles on its practical application. In order to amend that, this paper focuses on the research of a deep learning-based lip recognition application system, i.e., the design and development of a speech correction system for the hearing impaired, which aims to lay the foundation for the comprehensive implementation of automatic lip recognition technology in the future. First, we used a MobileNet lightweight network to extract spatial features from the original lip image; the extracted features are robust and fault-tolerant. Then, the gated recurrent unit (GRU) network was used to further extract the 2D image features and temporal features of the lip. To further improve the recognition rate, based on the GRU network, we incorporated an attention mechanism; the performance of this model is illustrated through a large number of experiments. Meanwhile, we constructed a lip similarity matching system to assist hearing-impaired people in learning and correcting their mouth shape with correct pronunciation. The experiments finally show that this system is highly feasible and effective.
随着深度学习和人工智能的发展,唇语识别在计算机视觉和人机交互领域的应用需求很高。特别是,利用自动唇语识别技术来提高听力障碍者在社交互动中的表现以及发音,是人工智能在医疗保健和康复领域最有前景的应用之一。唇语识别是指通过分析动态动作来识别说话者表达的内容。目前,唇语识别研究主要集中在算法和计算性能方面,但关于其实际应用的研究文章相对较少。为了弥补这一不足,本文重点研究基于深度学习的唇语识别应用系统,即针对听力障碍者的语音矫正系统的设计与开发,旨在为未来自动唇语识别技术的全面应用奠定基础。首先,我们使用MobileNet轻量级网络从原始唇图像中提取空间特征;提取的特征具有鲁棒性和容错性。然后,使用门控循环单元(GRU)网络进一步提取唇部的二维图像特征和时间特征。为了进一步提高识别率,基于GRU网络,我们引入了注意力机制;通过大量实验说明了该模型的性能。同时,我们构建了一个唇相似度匹配系统,以帮助听力障碍者学习并纠正他们的口型以发出正确的发音。实验最终表明该系统具有高度的可行性和有效性。