Driver Drowsiness Prediction Using CNN-LSTM Model Based on Facial Expression and Eye Movement
Keywords:
CNN-LSTM, Driver Drowsiness, MediaPipe FaceMesh, ResNet50V, Eye Aspect Ratio.Abstract
Driver fatigue and drowsiness represent primary institutional catalysts for fatal highway traffic anomalies worldwide. This comprehensive investigation introduces an adaptive, multi-task deep learning architecture merging Convolutional Neural Networks and Long Short-Term Memory configurations to dynamically evaluate driver states through localized facial expressions and non-invasive ocular metrics. Utilizing MediaPipe FaceMesh, the framework maps 468 distinct landmark parameters under fluctuating illumination constraints to monitor regional variations across the eyes and mouth. The quantitative metrics are extracted by formulating real-time computations of the Eye Aspect Ratio and Mouth Aspect Ratio. Spatiotemporal feature representation is accomplished using a pre-trained ResNet50V2 feature extractor integrated with a 128-unit recurrent LSTM layer to process sequences across a 20-frame context window. The multi-branch dense layer concurrently outputs predictive status conditions for drowsiness, yawning frequency, and categorical facial expressions. System validation metrics indicate an overall classification accuracy of 88% for structural drowsiness tracking and 90% for yawning anomalies. This system is a “proof of concept” designed to be implemented for drivers to reduce traffic accidents caused by driver fatigue.
Downloads
References
Agiomavritis, F., & Karanasiou, I. (2026). Tactical-grade wearables and authentication biometrics. Sensors, 26(3), 759. https://doi.org/10.3390/s26030759
AL-Quraishi, M. S., Azhar Ali, S. S., AL-Qurishi, M., Tang, T. B., & Elferik, S. (2024). Technologies for detecting and monitoring drivers’ states: A systematic review. Heliyon, 10(20), e39592. https://doi.org/10.1016/j.heliyon.2024.e39592
Asperti, A., & Filippini, D. (2023). Deep learning for head pose estimation: A survey. SN Computer Science, 4(4), 349. https://doi.org/10.1007/s42979-023-01796-z
Balazadeh Meresht, N., Moghadasi, S., Munshi, S., Shahbakhti, M., & McTaggart-Cowan, G. (2023). Advances in vehicle and powertrain efficiency of long-haul commercial vehicles: A review. Energies, 16(19), 6809. https://doi.org/10.3390/en16196809
Bhatta, Y. R., et al. (2025). Road traffic accidents in Prithvi Highway, Nepal: A spatiotemporal analysis. Journal of Engineering Technology and Planning, 6(1), 44–54. https://doi.org/10.3126/joetp.v6i1.87815
Cao, W., Lu, P., & Cao, W. (2024). Multimodal gesture recognition with spatio-temporal features fusion based on YOLOv5 and MediaPipe. International Journal of Pattern Recognition and Artificial Intelligence, 38(8). https://doi.org/10.1142/S0218001424550073
Dablain, D., Jacobson, K. N., Bellinger, C., Roberts, M., & Chawla, N. V. (2024). Understanding CNN fragility when learning with imbalanced data. Machine Learning, 113(7), 4785–4810. https://doi.org/10.1007/s10994-023-06326-9
Darwich, M., & Bayoumi, M. (2024). Video quality adaptation using CNN and RNN models for cost-effective and scalable video streaming services. Cluster Computing, 27(5), 6355–6375. https://doi.org/10.1007/s10586-024-04315-8
Dileep Kumar, M. J., Sukesh Rao, M., & Narendra, K. C. (2025). Multimodal emotion recognition: A comprehensive survey of datasets, methods, and applications. IEEE Access, 13, 201067–201097. https://doi.org/10.1109/ACCESS.2025.3636186
El-Nabi, S. A., El-Shafai, W., El-Rabaie, E.-S. M., Ramadan, K. F., Abd El-Samie, F. E., & Mohsen, S. (2024). Machine learning and deep learning techniques for driver fatigue and drowsiness detection: A review. Multimedia Tools and Applications, 83(3), 9441–9477. https://doi.org/10.1007/s11042-023-15054-0
Fife, S. T., & Gossner, J. D. (2024). Deductive qualitative analysis: Evaluating, expanding, and refining theory. International Journal of Qualitative Methods, 23. https://doi.org/10.1177/16094069241244856
Fu, S., Yang, Z., Ma, Y., Li, Z., Xu, L., & Zhou, H. (2024). Advancements in the intelligent detection of driver fatigue and distraction: A comprehensive review. Applied Sciences, 14(7), 3016. https://doi.org/10.3390/app14073016
Ghosh, K., Bellinger, C., Corizzo, R., Branco, P., Krawczyk, B., & Japkowicz, N. (2024). The class imbalance problem in deep learning. Machine Learning, 113(7), 4845–4901. https://doi.org/10.1007/s10994-022-06268-8
Jafari, F., Moradi, K., & Shafiee, Q. (2024). Shallow learning vs. deep learning in engineering applications (pp. 29–76). https://doi.org/10.1007/978-3-031-69499-8_2
John, C., Sahoo, J., Madhavan, M., & Mathew, O. K. (2023). Convolutional neural networks: A promising deep learning architecture for biological sequence analysis. Current Bioinformatics, 18(7), 537–558. https://doi.org/10.2174/1574893618666230320103421
Karim, R. U., Mahdi, S., Samin, A., Zereen, A. N., Abdullah-Al-Wadud, M., & Uddin, J. (2025). Optimizing stroke recognition with MediaPipe and machine learning: An explainable AI approach for facial landmark analysis. IEEE Access, 13, 32636–32660. https://doi.org/10.1109/ACCESS.2025.3550577
Khanal, S. R., Paulino, D., Sampaio, J., Barroso, J., Reis, A., & Filipe, V. (2022). A review on computer vision technology for physical exercise monitoring. Algorithms, 15(12), 444. https://doi.org/10.3390/a15120444
Li, J., et al. (2024). A review of computer vision-based monitoring approaches for construction workers’ work-related behaviors. IEEE Access, 12, 7134–7155. https://doi.org/10.1109/ACCESS.2024.3350773
Mani, Z. A., & Goniewicz, K. (2023). Transportation disaster trends and impacts in Western Asia: A comprehensive analysis from 2003 to 2023. Sustainability, 15(18), 13636. https://doi.org/10.3390/su151813636
Mao, D., Chen, Y., Wu, Y., Gilles, M., & Wong, A. (2024). Rethinking resource competition in multi-task learning: From shared parameters to shared representation. IEEE Access, 12, 128717–128728. https://doi.org/10.1109/ACCESS.2024.3429281
Montañez, J. J. F., & Alon, A. S. (2025). Enhancing face mask detection using ResNet-50 and MobileNetV2: A transfer learning application. Mindanao Journal of Science and Technology, 23(2), 311–334. https://doi.org/10.61310/mjst.v23i2.2502
Ramzan, M., Abid, A., Fayyaz, M., Alahmadi, T. J., Nobanee, H., & Rehman, A. (2024). A novel hybrid approach for driver drowsiness detection using a custom deep learning model. IEEE Access, 12, 126866–126884. https://doi.org/10.1109/ACCESS.2024.3438617
Selvaraju, V., et al. (2022). Continuous monitoring of vital signs using cameras: A systematic review. Sensors, 22(11), 4097. https://doi.org/10.3390/s22114097
Stounberg, J., Thomsen, T. B., Heredia, B. D., & Hüssy, K. (2022). Eyes and ears: A comparative approach linking the chemical composition of cod otoliths and eye lenses. Journal of Fish Biology, 101(4), 985–995. https://doi.org/10.1111/jfb.15159
Tasci, M. (2026). H-DrowsyNet: A dual-branch hybrid architecture based on Eye Aspect Ratio (EAR) and Convolutional Neural Networks (CNN) for driver fatigue detection. Akıllı Ulaşım Sistemleri ve Uygulamaları Dergisi, 9(1), 1–21. https://doi.org/10.51513/jitsa.1844823
Tsai, P.-S., Wu, T.-F., & Chen, W.-H. (2026). Generalized vision-based coordinate extraction framework for EDA layout reports and PCB optical positioning. Processes, 14(2), 342. https://doi.org/10.3390/pr14020342
Tumminello, M. L., Macioszek, E., & Granà, A. (2025). Emerging cutting-edge technologies and applications for safer, sustainable, and intelligent road systems in smart cities: A review. Applied Sciences, 15(21), 11583. https://doi.org/10.3390/app152111583
Visconti, P., Rausa, G., Del-Valle-Soto, C., Velázquez, R., Cafagna, D., & De Fazio, R. (2025). Innovative driver monitoring systems and on-board-vehicle devices in a smart-road scenario based on the Internet of Vehicle paradigm: A literature and commercial solutions overview. Sensors, 25(2), 562. https://doi.org/10.3390/s25020562
Wosiak, A., Sumiński, M., & Żykwińska, K. (2025). Impact of temporal window shift on EEG-based machine learning models for cognitive fatigue detection. Algorithms, 18(10), 629. https://doi.org/10.3390/a18100629
Xiao, L., Cao, Y., Gai, Y., Khezri, E., Liu, J., & Yang, M. (2023). Recognizing sports activities from video frames using deformable convolution and adaptive multiscale features. Journal of Cloud Computing, 12(1), 167. https://doi.org/10.1186/s13677-023-00552-1
Zaky, M. H., Shoorangiz, R., Poudel, G. R., Yang, L., Innes, C. R. H., & Jones, R. D. (2023). Increased cerebral activity during microsleeps reflects an unconscious drive to re-establish consciousness. International Journal of Psychophysiology, 189, 57–65. https://doi.org/10.1016/j.ijpsycho.2023.05.349
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 ARMADA : Jurnal Penelitian Multidisiplin

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.





