Authors:
Vijay Singh Rana,Ankush Joshi,Kamal Kant Verma,DOI NO:
https://doi.org/10.26782/jmcms.2026.08.00010Keywords:
HAR,ResNet50,Temporal Convolution,BiLSTM,Multimodal Fusion,RGBD,Florence 3d Dataset,Abstract
In recent years, researchers have become interested in human action recognition because of the broad spectrum of its applications in surveillance, healthcare monitoring, and human-computer interaction. A highly robust, efficient, and innovative multimodal system for recognizing human actions is proposed in this study based on the Florence 3D Action dataset. The proposed system leverages the RGB, Depth, and Skeleton multi-modalities. For the case of RGB and Depth video sequences, spatial information is extracted by a pre-trained ResNet50 model, and then the extracted high-level visual features are fed to a Temporal Convolutional Network (TCN) to capture temporal patterns at a large scale. Color and depth video sequences are processed in parallel. Temporal information is also captured in the skeleton data, which is fed to a Bi-directional Long Short-Term Memory (BiLSTM) model after features are extracted in the form of joint angles, velocities, and accelerations, which describe the motion of human actions. For the case of the three modalities, the results are fused at the decision level using a weighted sum for the final prediction. The results show that the proposed system has an impressive recognition accuracy of 96.8% on the Florence 3D dataset, with stable convergence, low validation loss and overfitting, and a strong generalization capability, which meets the requirements for use in real-time, resource-constrained systems.Refference:
I. Al-qaness, Mohammed AA, et al. "TCN-inception: temporal convolutional network and inception modules for sensor-based human activity recognition." Future Generation Computer Systems 160 (2024): 375-388. 10.1016/j.future.2024.06.016.
II. Anirudh, Rushil, et al. "Elastic functional coding of human actions: From vector-fields to latent variables." Proceedings of the IEEE conference on computer vision and pattern recognition. 2015. 10.1109/CVPR.2015.7298934.
III. Bai, Ruwen, et al. "Hierarchical graph convolutional skeleton transformer for action recognition." 2022 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 2022. 10.1109/ICME52920.2022.9859781.
IV. Bai, Shaojie, J. Zico Kolter, and Vladlen Koltun. "An empirical evaluation of generic convolutional and recurrent networks for sequence modeling." arXiv preprint arXiv:1803.01271 (2018). 10.48550/arXiv.1803.01271.
V. Batool, Mouazma, et al. "Multimodal human action recognition framework using an improved CNNGRU classifier." IEEE Access 12 (2024): 158388-158406. 10.1109/ACCESS.2024.3481631.
VI. Bijrothiya, Sadhna, and Vaibhav Soni. "An architecture for human activity recognition using TCN-BI-LSTM HAR based on wearable sensor." Procedia Computer Science 260 (2025): 805-813. 10.1016/j.procs.2025.03.261.
VII. Biswas, Sougatamoy, Anup Nandy, and Asim Kumar Naskar. "Lightweight multimodal feature fusion and spatiotemporal learning for human action recognition on edge devices." IEEE Transactions on Emerging Topics in Computational Intelligence (2025). 10.1109/TETCI.2025.3616054.
VIII. Chen, Junjie, et al. "A data augmentation method for skeleton-based action recognition with relative features." Applied Sciences 11.23 (2021): 11481. 10.3390/app112311481.
IX. Dhiman, Chhavi, Dinesh Kumar Vishwakarma, and Paras Agarwal. "Part-wise spatio-temporal attention driven CNN-based 3D human action recognition." ACM Transactions on Multimidia Computing Communications and Applications 17.3 (2021): 1-24. 10.1145/3441628.
X. Farha, Yazan Abu, and Jurgen Gall. "Ms-tcn: Multi-stage temporal convolutional network for action segmentation." Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019. 10.1109/CVPR.2019.00369.
XI. Funahashi, Ken-ichi, and Yuichi Nakamura. "Approximation of dynamical systems by continuous time recurrent neural networks." Neural networks 6.6 (1993): 801-806. 10.1016/S0893-6080(05)80125-X.
XII. Hidayanto, Nur Awal, Adhi Prahara, and Riky Dwi Puriyanto. "Hierarchical long short-term memory for action recognition based on 3D skeleton joints from Kinect sensor." Jurnal Informatika Ahmad Dahlan 15.1 (2021): 17-27. 10.26555/jifo.v15i1.a20106.
XIII. Hochreiter, Sepp, and Jürgen Schmidhuber. "Long short-term memory." Neural computation 9.8 (1997): 1735-1780. 10.1162/neco.1997.9.8.1735.
XIV. Hou, Jingxuan, Tae Soo Kim, and Austin Reiter. "Train, diagnose and fix: Interpretable approach for fine-grained action recognition." arXiv preprint arXiv:1711.08502 (2017). 10.48550/arXiv.1711.08502.
XV. Lea, Colin, et al. "Temporal convolutional networks for action segmentation and detection." 2017 IEEE conference on computer vision and pattern recognition (CVPR). IEEE, 2017. 10.48550/arXiv.1611.05267.
XVI. Liu, Chengming, et al. "Angle information assisting skeleton-based actions recognition." PeerJ Computer Science 10 (2024): e2523. 10.7717/peerj-cs.2523.
XVII. Murad, Abdulmajid, and Jae-Young Pyun. "Deep recurrent neural networks for human activity recognition." Sensors 17.11 (2017): 2556, 10.3390/s17112556.
XVIII. Musallam, Mohamed Adel, et al. "Temporal 3d human pose estimation for action recognition from arbitrary viewpoints." 2019 international conference on computational science and computational intelligence (csci). IEEE, 2019. 10.1109/CSCI49370.2019.00052.
XIX. Oord, Aaron van den, et al. "Wavenet: A generative model for raw audio." arXiv preprint arXiv:1609.03499 (2016). 10.48550/arXiv.1609.03499.
XX. Prasad Raghava, "Fusion of RGB and Skeletal Data Using Gated Features for Human Action Recognition”. International Journal of Food and Nutritional Sciences 11.6(2022): 1-6. 10.1007/s12652-019-01239-9.
XXI. Rahayu, Endang Sri, et al. "Human activity classification using deep learning based on 3D motion feature." Machine Learning with Applications 12 (2023): 100461. .1016/j.mlwa.2023.100461.
XXII. Rana, Vijay Singh, Ankush Joshi, and Kamal Kant Verma. "A Unified Deep Learning Approach for Human Activity Recognition with RGB-D Data." 2025 International Conference on Innovations and Emerging Technologies In AI & Communication Systems (IETACS). IEEE, 2025. 10.1109/IETACS68750.2025.11385548.
XXIII. Sanchez-Caballero, Adrian, David Fuentes-Jimenez, and Cristina Losada-Gutiérrez. "Real-time human action recognition using raw depth video-based recurrent neural networks." Multimedia Tools and Applications 82.11 (2023): 16213-16235. 10.1109/ACCESS.2024.3481631.
XXIV. Seidenari, Lorenzo, et al. "Recognizing actions from depth cameras as weakly aligned multi-part bag-of-poses." Proceedings of the IEEE conference on computer vision and pattern recognition workshops. 2013. 10.1109/CVPRW.2013.77.
XXV. Shaikh, Muhammad Bilal, et al. "From CNNs to transformers in multimodal human action recognition: A survey." ACM Transactions on Multimedia Computing, Communications and Applications 20.8 (2024): 1-24. 10.1145/3664815.
XXVI. Shotton, Jamie, et al. "Real-time human pose recognition in parts from single depth images." CVPR 2011. Ieee, 2011. 10.1109/CVPR.2011.5995316.
XXVII. Siami-Namini, Sima, Neda Tavakoli, and Akbar Siami Namin. "The performance of LSTM and BiLSTM in forecasting time series." 2019 IEEE International conference on big data (Big Data). IEEE, 2019. 10.1109/BigData47090.2019.9005997.
XXVIII. Vemulapalli, Raviteja, Felipe Arrate, and Rama Chellappa. "Human action recognition by representing 3d skeletons as points in a lie group." Proceedings of the IEEE conference on computer vision and pattern recognition. 2014. 10.1109/ICALIP.2016.7846646.
XXIX. Verma, Kamal Kant, and Brij Mohan Singh. "Deep Multi-Model Fusion for Human Activity Recognition Using Evolutionary Algorithms." International Journal of Interactive Multimedia and Artificial Intelligence 7.2 (2021): 44-58. 10.9781/ijimai.2021.08.008.
XXX. Verma, Kamal Kant, Brij Mohan Singh, and Amit Dixit. "A review of supervised and unsupervised machine learning techniques for suspicious behavior recognition in intelligent surveillance system." International Journal of Information Technology 14.1 (2022): 397-410. 10.1007/s41870-019-00364-0.
XXXI. Verma, Kamal Kant, et al. "Two-stage human activity recognition using 2D-ConvNet." IJIMAI 6.2 (2020): 125-135. 10.9781/ijimai.2020.04.002.
XXXII. Vijay Singh Rana, Ankush Joshi, Kamal Kant Verma, "Lightweight 3DCNN-BiLSTM Model for Human Activity Recognition using Fusion of RGBD Video Sequences", International Journal of Information Technology and Computer Science(IJITCS), Vol.17, No.6, pp.176-193. 2025. 10.5815/ijitcs.2025.06.10
XXXIII. Wang, Jiang, et al. "Mining actionlet ensemble for action recognition with depth cameras." 2012 IEEE conference on computer vision and pattern recognition. IEEE, 2012. 10.1109/CVPR.2012.6247813.
XXXIV. Wei, Lianglei, et al. "A novel 3d human action recognition framework for video content analysis." International Conference on Multimedia Modeling. Cham: Springer International Publishing, 2018. 10.1007/978-3-319-73603-7_4.
XXXV. Wei, Xiong, and Zifan Wang. "TCN-attention-HAR: Human activity recognition based on attention mechanism time convolutional network." Scientific Reports 14.1 (2024): 7414. 10.1038/s41598-024-57912-3.
XXXVI. Xie, Dongwei, et al. "MAF-Net: A multimodal data fusion approach for human action recognition." PloS one 20.4 (2025): e0319656. 10.1371/journal.pone.0319656.
XXXVII. Zhan, Tianyu. "Research on Action Recognition Algorithm Based on Multimodal Data Fusion." 2025 IEEE 5th International Conference on Electronic Technology, Communication and Information (ICETCI). IEEE, 2025. 10.1109/ICETCI64844.2025.11084160.
XXXVIII. Zhang, Yumin, and Yanyong Wang. "A comprehensive survey on RGB-D-based human action recognition: algorithms, datasets, and popular applications." EURASIP Journal on Image and Video Processing 2025.1 (2025): 15. 10.1186/s13640-025-00677-0.

