Back close

Hand Gesture Segmentation using UNet-VAE Deep Learning Architecture

Publication Type : Conference Proceedings

Publisher : IEEE

Source : 2025 5th International Conference on Intelligent Technologies (CONIT)

Url : https://doi.org/10.1109/conit65521.2025.11167307

Campus : Coimbatore

School : School of Artificial Intelligence - Coimbatore

Year : 2025

Abstract : Hand segmentation is the process of recognizing and separating the hand from the background. Segmenting the hand from the background plays an important role in several applications in computer vision tasks like gesture recognition, human-machine interaction and much more. This manuscript is mainly focused on segmenting the hand using a UNet architecture integrated with variational auto encoder (VAE). The VAE is integrated in the bottle neck layer of the UNet model, which can learn the distribution of the input data and produce better segmentation mask compared to traditional UNet architecture. Several open source datasets such as the EgoHands dataset, Roboflow dataset and EgoYouTubeHands (EYTH) dataset are employed to evaluate model’s performance. A combination of KL divergence loss and reconstruction losses are used to quantify the error between the actual and the predicted. The evaluation metrics such as IoU and Dice score are calculated to measure how well the prediction aligns with the actual labels. The proposed architecture achieved a Dice score of 0.89 and an IoU of 0.95 on the EgoHands dataset, 0.91 and 0.90 on the EgoYouTubeHands dataset, and 0.94 and 0.88 on the Roboflow dataset respectively. In comparison, the UNet model obtained a Dice score of 0.74 and an IoU of 0.66 on EgoHands dataset, 0.72 and 0.63 on EgoYouTubeHands dataset, and 0.90 and 0.85 on Roboflow dataset. Similarly, UNet with Attention achieved a Dice score of 0.82 and an IoU of 0.73 on EgoHands dataset, 0.84 and 0.73 on EgoYouTubeHands dataset, and 0.92 and 0.87 on Roboflow dataset. These results demonstrates that the proposed model outperforms both UNet and UNet with Attention across all three datasets.

Cite this Research Publication : Meghaa Elangovan, Mithun Kumar Kar, Hand Gesture Segmentation using UNet-VAE Deep Learning Architecture, 2025 5th International Conference on Intelligent Technologies (CONIT), IEEE, 2025, https://doi.org/10.1109/conit65521.2025.11167307

Admissions Apply Now