Publication Type : Conference Proceedings
Publisher : IEEE
Source : 2022 International Conference on Innovative Trends in Information Technology (ICITIIT)
Url : https://doi.org/10.1109/icitiit54346.2022.9744134
Campus : Amritapuri
School : School of Computing
Year : 2022
Abstract : In machine learning and data mining, String Kernels combined with classifiers like Support Vector Machines (SVM) show state-of-the-art results for tasks such as text classification. Traditional pairwise comparisons of strings on large datasets are computationally expensive and result in quadratic runtimes. This work compares the performance of various String Kernels and similarity measures on the document classification task. We compare different String Kernels such as Spectrum Kernel, String Subsequence Kernel, Weighted Degree Kernel, and Distance Substitution Kernel in this paper for classifying text documents. A detailed comparative study of these Kernel techniques on real-life document corpus such as Reuters-21578 shows different insights when used with and without other feature extraction techniques. The results indicate that string similarity measures give the best performance when run over the entire corpus but for small and medium-sized datasets. The complexity increases with an increase in the size of the dataset.
Cite this Research Publication : Nikhil V. Chandran, Asharaf S., Anoop V. S., String Kernels for Document Classification: A Comparative Study, 2022 International Conference on Innovative Trends in Information Technology (ICITIIT), IEEE, 2022, https://doi.org/10.1109/icitiit54346.2022.9744134