Publication Type : Conference Proceedings
Publisher : IEEE
Source : 2025 5th International Conference on Ubiquitous Computing and Intelligent Information Systems (ICUIS)
Url : https://doi.org/10.1109/icuis67429.2025.11380688
Campus : Bengaluru
School : School of Engineering
Department : Electronics and Communication
Year : 2025
Abstract : One of the major issues in health research is that the patient data necessary to develop good predictive models is often not available because of the privacy laws and regulations. In this research, Conditional Tabular Generative Adversarial Networks (CTGAN) are looked at for the production of synthetic health data that is privacy-preserving and their use in diabetes risk prediction is assessed. The Pima Indians Diabetes dataset was used to train Logistic Regression, Random Forest, and Support Vector Machine classifiers on both the original and CTGAN-generated synthetic datasets. The model performance was evaluated through accuracy, precision, recall, F1-score, and AUC metrics. The outcome showed that the Random Forest was the best with an accuracy of 82% on original and 63% on synthetic datasets, proving that synthetic data had a large share of the predictive power while being privacy-safe. The decrease in performance was not too much but still allowed the data to keep the main statistical relationships, thus confirming CTGAN as a suitable tool for privacy-conscious healthcare analytics. The research presents synthetic data as a solution to the problem of medical AI being limited to non-confidential areas by the need to be secure, ethical, and reproducible.
Cite this Research Publication : Prabina Subedi, Sri Ramya Divakarla, Manaswini Emani, Sreeja Kochuvila, Sunitha R., Synthetic Data for Health Risk Prediction: An Evaluation using Machine Learning, 2025 5th International Conference on Ubiquitous Computing and Intelligent Information Systems (ICUIS), IEEE, 2025, https://doi.org/10.1109/icuis67429.2025.11380688