Publication Type : Journal Article
Publisher : Elsevier BV
Source : Computers and Electrical Engineering
Url : https://doi.org/10.1016/j.compeleceng.2026.111004
Keywords : Security, Malicious URLs, Machine learning, Class balancing, Phishing attack, ISCX-URL2016
Campus : Amaravati
School : School of Computing
Year : 2026
Abstract : With the advancement of technology, cyberattacks on Internet-based services such as email, e-commerce, social networking, and electronic healthcare are increasing. Since many of these services are accessed through URLs, they have become a primary source for cyberattacks, including phishing and malware. Anti-Phishing Working Group (APWG) reported nearly 1 million phishing attacks in the first quarter of 2025. Early detection of malicious URLs is therefore critical to preventing these threats. Therefore, an efficient detection of malicious URLs is an emerging research problem. However, most ML/DL-based studies focus on overall model accuracy and tend to be biased towards majority classes in imbalanced datasets. In this paper, we propose a machine learning-based malicious URL detection framework specifically designed for imbalanced datasets. We use the ISCX-URL2016 dataset to evaluate model performance across multiple ML algorithms and classbalancing techniques. Our proposed framework, combining the LightGBM classifier with ADASYN oversampling, achieves 99.76% accuracy in multi-class and 99.92% in binary classification. Notably, it shows a 5.93% improvement in detecting phishing URLs, a minority class in the dataset, over existing models. A significant achievement of our approach is its uniform performance across all classes, effectively reducing bias towards majority classes, while existing models fail to achieve it, particularly minority classes. We also validated the proposed model using recent datasets. We further evaluate the framework using various feature selection techniques, demonstrating its effectiveness with fewer features. Additionally, we perform statistical significance testing to validate the reliability of our model, confirming its suitability for real-world applications.
Cite this Research Publication : K.G. Raghavendra Narayan, Srijanee Mookherji, Vanga Odelu, Rajendra Prasath, A framework for handling class imbalance in malicious URL dataset, Computers and Electrical Engineering, Elsevier BV, 2026, https://doi.org/10.1016/j.compeleceng.2026.111004