Computer Vision Techniques in Violence Detection: An Overview
DOI:
https://doi.org/10.61704/pr.567Keywords:
Computer Vision, Violence Detection Techniques, CNN, RNN, LSTM, YOLOAbstract
The current research paper discusses the importance of computer vision systems in every facet of life including security, health, commerce, and human-machine interaction. The paper shows that there are numerous computer vision methods of tracking human behavior and detecting violence in video. These techniques are convolutional neural networks (CNNs) to extract spatial features, deep time networks, including RNNs and LSTMs, to analyze the time sequence of events, transformer-based models, including ViViT and TimeSformer, to analyze long sequences of spatial and temporal data, and active optical flow methods to detect sudden changes between images. It is also used in the detection of violence cases, which is based on the posture of people and YOLO or object detection. Moreover, audio analysis and multimodal fusion are used to improve the algorithm of violence detection and supervisory self-learning, advanced spatiotemporal analysis, and matrix analysis are used to present more information about individual behavior. The pros and cons of computer vision-based violence detection algorithms are also compared.
References
Abdullah, D. B., & Alnuaimy, M. R. (2022). Real-time face tracking for service-robot. Technium: Romanian Journal of Applied Sciences and Technology, 4(9), 47–52. https://doi.org/10.47577/technium.v4i9.7330
Ahmad, B., Khan, M., & Sajjad, M. (2025). Gated fusion networks for multi-modal violence detection. AI, 6(10), Article 259. https://doi.org/10.3390/ai6100259
Al-Hayali, A. I. B. (2018). Hiding and retrieving encrypted data using LSB in an image based on RBF network. AL-Rafidain Journal of Computer Sciences and Mathematics, 1(1), 199-209 https://stats.uomosul.edu.iq/index.php/csmj/article/view/37053/36843
Anwar, S., Ullah, A., Rocha, Á., & Sousa, M. J. (Eds.). (2023). Proceedings of International Conference on Information Technology and Applications: ICITA 2022 (Lecture Notes in Networks and Systems, Vol. 614). Springer. https://doi.org/10.1007/978-981-19-9331-2
Alisetti, S. V., Kalagi, V., & Krishnagopal, S. (2025). Beyond Attention: Learning Spatio-Temporal Dynamics with Emergent Interpretable Topologies. arXiv preprint arXiv:2506.00770.
Arenas, R., Méndez, R., Pedraza, L., Castelló, E., & Flores, J. (2025). Human pose estimation solutions: a low cost tool for increasing natural interaction in virtual television sets. Universal Access in the Information Society, 24(4), 3257-3270.
Bakhshi, A., García-Gómez, J., Gil-Pita, R., & Chalup, S. (2023). Violence detection in real-life audio signals using lightweight deep neural networks. Procedia Computer Science, 222, 244-251.
Chang, V., Eniola, R. O., Golightly, L., & Xu, Q. A. (2023). An exploration into human–computer interaction: Hand gesture recognition management in a challenging environment. SN Computer Science, 4(5), 441.
Calderon-Vilca, D., Cuadros-Ramos, K., Valcarcel-Ascencios, S., & Aguilar-Alonso, I. (2025). Hybrid CNN Xception and Long Short-Term Memory Model for the Detection of Interpersonal Violence in Videos. International Arab Journal of Information Technology (IAJIT), 22(5).
Du, S., & Ikenaga, T. (2025). Human pose analysis: Deep learning meets human kinematics in video. Springer. https://doi.org/10.1007/978-981-97-9334-1
Dilek, E., & Dener, M. (2025). An overview of transformers for video anomaly detection. Neural Computing and Applications, 1-33. https://link.springer.com/article/10.1007/s00521-025-11218-1
Durães, D., Veloso, B., & Novais, P. (2023). Violence detection in audio: evaluating the effectiveness of deep learning models and data augmentation. DOI: https://doi.org/10.9781/ijimai.2023.08.007
Elhoseny, M., Lydia, E. L., Sree, S. R., Akhmetshin, E., & Shankar, K. (2025). Lightweight Convolutional Neural Network Based Computer Vision Model for Human Behaviour Analysis on Consumer Internet of Things Devices. IEEE Transactions on Consumer Electronics. https://ieeexplore.ieee.org/abstract/document/10980367/
Elhanashi, A., Dini, P., Saponara, S., & Zheng, Q. (2024). TeleStroke: real-time stroke detection with federated learning and YOLOv8 on edge devices. Journal of Real-Time Image Processing, 21(4), 121. https://link.springer.com/article/10.1007/s11554-024-01500-1
Elek, R., Károly, A., Haidegger, T., & Galambos, P. (2020, January). Towards optical flow ego-motion compensation for moving object segmentation. In Proceedings of the International Conference on Robotics, Computer Vision and Intelligent Systems. https://doi.org/10.5220/0010136301140120
Faouzi, J., & Colliot, O. (2023). Classic machine learning methods. Machine learning for brain disorders, 25-75. https://doi.org/10.1007/978-1-0716-3195-9_2
Farnebäck, G. (2003, June). Two-frame motion estimation based on polynomial expansion. In Scandinavian conference on Image analysis (pp. 363-370). Berlin, Heidelberg: Springer Berlin Heidelberg. https://doi.org/10.1007/3-540-45103-X_50
Gaya-Morey, F. X., Manresa-Yee, C., & Buades-Rubio, J. M. (2024). Deep learning for computer vision-based activity recognition and fall detection of the elderly: a systematic review: GM F. Xavier et al. Applied Intelligence, 54(19), 8982-9007.
Garcia-Cobo, G., & SanMiguel, J. C. (2023). Human skeletons and change detection for efficient violence detection in surveillance videos. Computer Vision and Image Understanding, 233, 103739.
Guo, Q., Tan, Q., Peng, Y., Xiao, L., Liu, M., & Shi, B. (2025). Model-enhanced spatial-temporal attention networks for traffic density prediction. Complex & Intelligent Systems, 11(1), 19.
Huang, H., & Jiang, Q. (2025). IDG-ViolenceNet: A Video Violence Detection Model Integrating Identity-Aware Graphs and 3D-CNN. Sensors, 25(20), 6272.
Hanson, A., Pnvr, K., Krishnagopal, S., & Davis, L. (2018). Bidirectional convolutional lstm for the detection of violence in videos. In Proceedings of the European conference on computer vision (ECCV) workshops.
Halder, R., & Chatterjee, R. (2020). CNN-BiLSTM Model for Violence Detection in Smart Surveillance: R. Halder, R. Chatterjee. SN Computer science, 1(4), 201.
Hussain, M. (2024). Yolov1 to v8: Unveiling each variant–a comprehensive review of yolo. IEEE access, 12, 42816-42833.
Hermens, F. (2024). Automatic object detection for behavioural research using YOLOv8. Behavior research methods, 56(7), 7307-7330.
Haque, M. A. H. M. U. D. U. L. (2022). Detection and classification of sensitive audio-visual content for automated film censorship and rating.
Ibrahim, R. T., & Ramo, F. M. (2023). Hybrid intelligent technique with deep learning to classify personality traits. International Journal of Computing and Digital Systems, 13(1), 231-244.
Jin, W., Zhu, L., & Sun, J. (2025). Aligning First, Then Fusing: A novel weakly supervised multimodal violence detection method. Knowledge-Based Systems, 322, 113709.
Jain, D. M., Bora, M. S., Chandnani, S., Grover, S., & Sadwal, S. (2023). Comparison of VGG-16, VGG-19, and ResNet-101 CNN models for the purpose of suspicious activity detection. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 121-130.
Kaur, G., & Singh, S. (2025). An ensemble based approach for violence detection in videos using deep transfer learning. Multimedia Tools and Applications, 84(12), 11001-11025.
Dong-Hyun, K., & Gratchev, I. (2021). Application of optical flow technique and photogrammetry for rockfall dynamics: A case study on a field test. Remote Sensing, 13(20), 4124.
Lohithashva, B. H., Aradhya, V. N., & Guru, D. S. (2020). Violent Video Event Detection Based on Integrated LBP and GLCM Texture Features. Revue d'Intelligence Artificielle, 34(2). https://doi.org/10.18280/ria.340208
Mahmoodi, J., & Nezamabadi-pour, H. (2024). A spatio-temporal model for violence detection based on spatial and temporal attention modules and 2D CNNs. Pattern Analysis and Applications, 27(2), 46.
Maji, D., Nagori, S., Mathew, M., & Poddar, D. (2022). Yolo-pose: Enhancing yolo for multi person pose estimation using object keypoint similarity loss. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 2637-2646). https://doi.org/10.48550/arXiv.2204.06806
Negre, P., Alonso, R. S., González-Briones, A., Prieto, J., & Rodríguez-González, S. (2024). Literature review of deep-learning-based detection of violence in video. Sensors, 24(12), 4016.
Nardelli, P., & Comminiello, D. (2024). Josenet: A joint stream embedding network for violence detection in surveillance videos. arXiv preprint arXiv:2405.02961. https://doi.org/10.48550/arXiv.2405.02961
Omarov, B., Narynov, S., Zhumanov, Z., Gumar, A., & Khassanova, M. (2022). A skeleton-based approach for campus violence detection. Computers, Materials & Continua, 72(1).
Paolanti, M., Pietrini, R., Mancini, A., Frontoni, E., & Zingaretti, P. (2020). Deep understanding of shopper behaviours and interactions using RGB-D vision. Machine Vision and Applications, 31(7), 66.
Park, J. H., Mahmoud, M., & Kang, H. S. (2024). Conv3D-based video violence detection network using optical flow and RGB data. Sensors, 24(2), 317. https://doi.org/10.3390/s24020317
Raouf, N. N., & Aldabbagh, M. T. (2024). Developing a Third-Party API to Enhance Image Documents at the University of Mosul Data Center. International Journal of Computing, 23(4), 692-701.
Rendón-Segador, F. J., Álvarez-García, J. A., Salazar-González, J. L., & Tommasi, T. (2023). Crimenet: Neural structured learning using vision transformer for violence detection. Neural networks, 161, 318-329.
Renugadevi, P., & Kalpana, S. (2025). A self-supervised lightweight CNN-LSTM network for violence detection in low-labelled surveillance video. International Journal of Scientific Research and Engineering Development, 8(4), 914–918. https://doi.org/10.5281/zenodo.16411318
Sebastian, R. A., Ehinger, K., & Miller, T. (2025). Do we need watchful eyes on our workers? Ethics of using computer vision for workplace surveillance. AI and Ethics, 5(4), 3557-3577.
Swapnil, A. L., Peris, M. D., Nihal, I. H., Khan, R., & Matin, M. A. (2024). Multimodal Deep Learning for Violence Detection: VGGish and MobileViT Integration with Knowledge Distillation on Jetson Nano. IEEE Open Journal of the Communications Society, 6, 2907-2925.
Singh, S., Dewangan, S., Krishna, G. S., Tyagi, V., Reddy, S., & Medi, P. R. (2022). Video vision transformers for violence detection. arXiv preprint arXiv:2209.03561. https://doi.org/10.48550/arXiv.2209.03561
Salehin, S., Rahman, S., Nur, M., Asif, A., Bin Harun, M., & Uddin, J. I. A. (2024). A Deep Learning Model for YOLOv9-based Human Abnormal Activity Detection: Violence and Non-Violence Classification. Iranian Journal of Electrical & Electronic Engineering, 20(4).
Senthilkumar, S., Agarwal, G., Shirish, A., & Kolte, S. (2025, February). Real Time Violence Detection System Using YOLOv7 and Deep Learning Techniques. In 2025 3rd International Conference on Intelligent Data Communication Technologies and Internet of Things (IDCIoT) (pp. 1447-1454). IEEE. https://doi.org/10.1109/IDCIOT64235.2025.10914712
Slade, S., Zhang, L., Yu, Y., & Lim, C. P. (2022). An evolving ensemble model of multi-stream convolutional neural networks for human action recognition in still images. Neural computing and applications, 34(11), 9205-9231.
Shin, J., Miah, A. S. M., Kaneko, Y., Hassan, N., Lee, H. S., & Jang, S. W. (2024). Multimodal attention-enhanced feature fusion-based weakly supervised anomaly violence detection. IEEE Open Journal of the Computer Society, 6, 129-140.
Varghese, E. B., Elzein, A., Yang, Y., & Qaraqe, M. (2025). A temporal–spatial deep learning framework leveraging dynamic 3D attention maps for violence detection. Neural Computing and Applications, 37(32), 26689-26709.
Wastupranata, L. M., Kong, S. G., & Wang, L. (2024). Deep Learning for Abnormal Human Behavior Detection in Surveillance Videos—A Survey. Electronics (2079-9292), 13(13).
Wu, P., Pan, C., Yan, Y., Pang, G., Yan, Q., Wang, P., & Zhang, Y. (2026). Deep learning for video anomaly detection: A review. IEEE Transactions on Neural Networks and Learning Systems.
Yu, J., Liu, J., Cheng, Y., Feng, R., & Zhang, Y. (2022, October). Modality-aware contrastive instance learning with self-distillation for weakly-supervised audio-visual violence detection. In Proceedings of the 30th ACM international conference on multimedia (pp. 6278-6287). https://doi.org/10.1145/3503161.3547868
Yildirim, S., Chimeumanu, M. S., & Rana, Z. A. (2023). The influence of micro-expressions on deception detection. Multimedia Tools and Applications, 82(19), 29115-29133. https://link.springer.com/article/10.1007/s11042-023-14551-6
Yang, Z., Du, H., Niyato, D., Wang, X., Zhou, Y., Feng, L., ... & Qiu, X. (2025). Revolutionizing wireless networks with self-supervised learning: A pathway to intelligent communications. IEEE Wireless Communications.
Yakar, İ., Kuçak, R. A., Bilgi, S., Ferhanoglu, O., & Akinci, T. C. (2025). A Hybrid Deep Learning and Optical Flow Framework for Monocular Capsule Endoscopy Localization. Electronics, 14(18), 3722.
Zhao, X., Wang, L., Zhang, Y., Han, X., Deveci, M., & Parmar, M. (2024). A review of convolutional neural networks in computer vision. Artificial Intelligence Review, 57(4), 99.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Fatimah A. Jasim, Ielaf O. AbdulMajjed

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Copyright © 2025 by the authors. This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License (CC BY-NC-ND 4.0). You may not alter or transform this work in any way without permission from the authors. Non-commercial use, distribution, and copying are permitted, provided that appropriate credit is given to the authors and Al-Hadba University.


