Back close

A framework for handling class imbalance in malicious URL dataset

Publication Type : Journal Article

Publisher : Elsevier BV

Source : Computers and Electrical Engineering

Url : https://doi.org/10.1016/j.compeleceng.2026.111004

Keywords : Security, Malicious URLs, Machine learning, Class balancing, Phishing attack, ISCX-URL2016

Campus : Amaravati

School : School of Computing

Year : 2026

Abstract : With the advancement of technology, cyberattacks on Internet-based services such as email, e-commerce, social networking, and electronic healthcare are increasing. Since many of these services are accessed through URLs, they have become a primary source for cyberattacks, including phishing and malware. Anti-Phishing Working Group (APWG) reported nearly 1 million phishing attacks in the first quarter of 2025. Early detection of malicious URLs is therefore critical to preventing these threats. Therefore, an efficient detection of malicious URLs is an emerging research problem. However, most ML/DL-based studies focus on overall model accuracy and tend to be biased towards majority classes in imbalanced datasets. In this paper, we propose a machine learning-based malicious URL detection framework specifically designed for imbalanced datasets. We use the ISCX-URL2016 dataset to evaluate model performance across multiple ML algorithms and classbalancing techniques. Our proposed framework, combining the LightGBM classifier with ADASYN oversampling, achieves 99.76% accuracy in multi-class and 99.92% in binary classification. Notably, it shows a 5.93% improvement in detecting phishing URLs, a minority class in the dataset, over existing models. A significant achievement of our approach is its uniform performance across all classes, effectively reducing bias towards majority classes, while existing models fail to achieve it, particularly minority classes. We also validated the proposed model using recent datasets. We further evaluate the framework using various feature selection techniques, demonstrating its effectiveness with fewer features. Additionally, we perform statistical significance testing to validate the reliability of our model, confirming its suitability for real-world applications.

Cite this Research Publication : K.G. Raghavendra Narayan, Srijanee Mookherji, Vanga Odelu, Rajendra Prasath, A framework for handling class imbalance in malicious URL dataset, Computers and Electrical Engineering, Elsevier BV, 2026, https://doi.org/10.1016/j.compeleceng.2026.111004

Admissions Apply Now