Driver Drowsiness Prediction Using CNN-LSTM Model Based on Facial Expression and Eye Movement

Authors

  • Annisa Aprilia Putri Sakri Electrical Engineering Study Program, Faculty of Electrical Engineering, Telkom University, Indonesia
  • Angga Rusdinar Electrical Engineering Study Program, Faculty of Electrical Engineering, Telkom University, Indonesia
  • Giva Andriana Mutiara Electrical Engineering Study Program, Faculty of Electrical Engineering, Telkom University, Indonesia
  • Periyadi Electrical Engineering Study Program, Faculty of Electrical Engineering, Telkom University, Indonesia

Keywords:

CNN-LSTM, Driver Drowsiness, MediaPipe FaceMesh, ResNet50V, Eye Aspect Ratio.

Abstract

Driver fatigue and drowsiness represent primary institutional catalysts for fatal highway traffic anomalies worldwide. This comprehensive investigation introduces an adaptive, multi-task deep learning architecture merging Convolutional Neural Networks and Long Short-Term Memory configurations to dynamically evaluate driver states through localized facial expressions and non-invasive ocular metrics. Utilizing MediaPipe FaceMesh, the framework maps 468 distinct landmark parameters under fluctuating illumination constraints to monitor regional variations across the eyes and mouth. The quantitative metrics are extracted by formulating real-time computations of the Eye Aspect Ratio and Mouth Aspect Ratio. Spatiotemporal feature representation is accomplished using a pre-trained ResNet50V2 feature extractor integrated with a 128-unit recurrent LSTM layer to process sequences across a 20-frame context window. The multi-branch dense layer concurrently outputs predictive status conditions for drowsiness, yawning frequency, and categorical facial expressions. System validation metrics indicate an overall classification accuracy of 88% for structural drowsiness tracking and 90% for yawning anomalies. This system is a “proof of concept” designed to be implemented for drivers to reduce traffic accidents caused by driver fatigue.

Downloads

Download data is not yet available.

References

Agiomavritis, F., & Karanasiou, I. (2026). Tactical-grade wearables and authentication biometrics. Sensors, 26(3), 759. https://doi.org/10.3390/s26030759

AL-Quraishi, M. S., Azhar Ali, S. S., AL-Qurishi, M., Tang, T. B., & Elferik, S. (2024). Technologies for detecting and monitoring drivers’ states: A systematic review. Heliyon, 10(20), e39592. https://doi.org/10.1016/j.heliyon.2024.e39592

Asperti, A., & Filippini, D. (2023). Deep learning for head pose estimation: A survey. SN Computer Science, 4(4), 349. https://doi.org/10.1007/s42979-023-01796-z

Balazadeh Meresht, N., Moghadasi, S., Munshi, S., Shahbakhti, M., & McTaggart-Cowan, G. (2023). Advances in vehicle and powertrain efficiency of long-haul commercial vehicles: A review. Energies, 16(19), 6809. https://doi.org/10.3390/en16196809

Bhatta, Y. R., et al. (2025). Road traffic accidents in Prithvi Highway, Nepal: A spatiotemporal analysis. Journal of Engineering Technology and Planning, 6(1), 44–54. https://doi.org/10.3126/joetp.v6i1.87815

Cao, W., Lu, P., & Cao, W. (2024). Multimodal gesture recognition with spatio-temporal features fusion based on YOLOv5 and MediaPipe. International Journal of Pattern Recognition and Artificial Intelligence, 38(8). https://doi.org/10.1142/S0218001424550073

Dablain, D., Jacobson, K. N., Bellinger, C., Roberts, M., & Chawla, N. V. (2024). Understanding CNN fragility when learning with imbalanced data. Machine Learning, 113(7), 4785–4810. https://doi.org/10.1007/s10994-023-06326-9

Darwich, M., & Bayoumi, M. (2024). Video quality adaptation using CNN and RNN models for cost-effective and scalable video streaming services. Cluster Computing, 27(5), 6355–6375. https://doi.org/10.1007/s10586-024-04315-8

Dileep Kumar, M. J., Sukesh Rao, M., & Narendra, K. C. (2025). Multimodal emotion recognition: A comprehensive survey of datasets, methods, and applications. IEEE Access, 13, 201067–201097. https://doi.org/10.1109/ACCESS.2025.3636186

El-Nabi, S. A., El-Shafai, W., El-Rabaie, E.-S. M., Ramadan, K. F., Abd El-Samie, F. E., & Mohsen, S. (2024). Machine learning and deep learning techniques for driver fatigue and drowsiness detection: A review. Multimedia Tools and Applications, 83(3), 9441–9477. https://doi.org/10.1007/s11042-023-15054-0

Fife, S. T., & Gossner, J. D. (2024). Deductive qualitative analysis: Evaluating, expanding, and refining theory. International Journal of Qualitative Methods, 23. https://doi.org/10.1177/16094069241244856

Fu, S., Yang, Z., Ma, Y., Li, Z., Xu, L., & Zhou, H. (2024). Advancements in the intelligent detection of driver fatigue and distraction: A comprehensive review. Applied Sciences, 14(7), 3016. https://doi.org/10.3390/app14073016

Ghosh, K., Bellinger, C., Corizzo, R., Branco, P., Krawczyk, B., & Japkowicz, N. (2024). The class imbalance problem in deep learning. Machine Learning, 113(7), 4845–4901. https://doi.org/10.1007/s10994-022-06268-8

Jafari, F., Moradi, K., & Shafiee, Q. (2024). Shallow learning vs. deep learning in engineering applications (pp. 29–76). https://doi.org/10.1007/978-3-031-69499-8_2

John, C., Sahoo, J., Madhavan, M., & Mathew, O. K. (2023). Convolutional neural networks: A promising deep learning architecture for biological sequence analysis. Current Bioinformatics, 18(7), 537–558. https://doi.org/10.2174/1574893618666230320103421

Karim, R. U., Mahdi, S., Samin, A., Zereen, A. N., Abdullah-Al-Wadud, M., & Uddin, J. (2025). Optimizing stroke recognition with MediaPipe and machine learning: An explainable AI approach for facial landmark analysis. IEEE Access, 13, 32636–32660. https://doi.org/10.1109/ACCESS.2025.3550577

Khanal, S. R., Paulino, D., Sampaio, J., Barroso, J., Reis, A., & Filipe, V. (2022). A review on computer vision technology for physical exercise monitoring. Algorithms, 15(12), 444. https://doi.org/10.3390/a15120444

Li, J., et al. (2024). A review of computer vision-based monitoring approaches for construction workers’ work-related behaviors. IEEE Access, 12, 7134–7155. https://doi.org/10.1109/ACCESS.2024.3350773

Mani, Z. A., & Goniewicz, K. (2023). Transportation disaster trends and impacts in Western Asia: A comprehensive analysis from 2003 to 2023. Sustainability, 15(18), 13636. https://doi.org/10.3390/su151813636

Mao, D., Chen, Y., Wu, Y., Gilles, M., & Wong, A. (2024). Rethinking resource competition in multi-task learning: From shared parameters to shared representation. IEEE Access, 12, 128717–128728. https://doi.org/10.1109/ACCESS.2024.3429281

Montañez, J. J. F., & Alon, A. S. (2025). Enhancing face mask detection using ResNet-50 and MobileNetV2: A transfer learning application. Mindanao Journal of Science and Technology, 23(2), 311–334. https://doi.org/10.61310/mjst.v23i2.2502

Ramzan, M., Abid, A., Fayyaz, M., Alahmadi, T. J., Nobanee, H., & Rehman, A. (2024). A novel hybrid approach for driver drowsiness detection using a custom deep learning model. IEEE Access, 12, 126866–126884. https://doi.org/10.1109/ACCESS.2024.3438617

Selvaraju, V., et al. (2022). Continuous monitoring of vital signs using cameras: A systematic review. Sensors, 22(11), 4097. https://doi.org/10.3390/s22114097

Stounberg, J., Thomsen, T. B., Heredia, B. D., & Hüssy, K. (2022). Eyes and ears: A comparative approach linking the chemical composition of cod otoliths and eye lenses. Journal of Fish Biology, 101(4), 985–995. https://doi.org/10.1111/jfb.15159

Tasci, M. (2026). H-DrowsyNet: A dual-branch hybrid architecture based on Eye Aspect Ratio (EAR) and Convolutional Neural Networks (CNN) for driver fatigue detection. Akıllı Ulaşım Sistemleri ve Uygulamaları Dergisi, 9(1), 1–21. https://doi.org/10.51513/jitsa.1844823

Tsai, P.-S., Wu, T.-F., & Chen, W.-H. (2026). Generalized vision-based coordinate extraction framework for EDA layout reports and PCB optical positioning. Processes, 14(2), 342. https://doi.org/10.3390/pr14020342

Tumminello, M. L., Macioszek, E., & Granà, A. (2025). Emerging cutting-edge technologies and applications for safer, sustainable, and intelligent road systems in smart cities: A review. Applied Sciences, 15(21), 11583. https://doi.org/10.3390/app152111583

Visconti, P., Rausa, G., Del-Valle-Soto, C., Velázquez, R., Cafagna, D., & De Fazio, R. (2025). Innovative driver monitoring systems and on-board-vehicle devices in a smart-road scenario based on the Internet of Vehicle paradigm: A literature and commercial solutions overview. Sensors, 25(2), 562. https://doi.org/10.3390/s25020562

Wosiak, A., Sumiński, M., & Żykwińska, K. (2025). Impact of temporal window shift on EEG-based machine learning models for cognitive fatigue detection. Algorithms, 18(10), 629. https://doi.org/10.3390/a18100629

Xiao, L., Cao, Y., Gai, Y., Khezri, E., Liu, J., & Yang, M. (2023). Recognizing sports activities from video frames using deformable convolution and adaptive multiscale features. Journal of Cloud Computing, 12(1), 167. https://doi.org/10.1186/s13677-023-00552-1

Zaky, M. H., Shoorangiz, R., Poudel, G. R., Yang, L., Innes, C. R. H., & Jones, R. D. (2023). Increased cerebral activity during microsleeps reflects an unconscious drive to re-establish consciousness. International Journal of Psychophysiology, 189, 57–65. https://doi.org/10.1016/j.ijpsycho.2023.05.349

Downloads

Published

2026-08-31

How to Cite

Annisa Aprilia Putri Sakri, Angga Rusdinar, Giva Andriana Mutiara, & Periyadi, P. (2026). Driver Drowsiness Prediction Using CNN-LSTM Model Based on Facial Expression and Eye Movement. ARMADA : Jurnal Penelitian Multidisiplin, 4(8), 3889–3903. Retrieved from https://ejournal.45mataram.ac.id/index.php/armada/article/view/3341