Publication Type : Conference Paper
Publisher : Springer International Publishing
Source : Advances in Intelligent Systems and Computing
Url : https://doi.org/10.1007/978-3-319-68385-0_11
Campus : Amritapuri
School : School of Computing
Year : 2017
Abstract : This paper proposes a framework which induces semantically rich concepts from probabilistically generated topics by a topic modeling algorithm. In this method an off-the-shelf tool has been used to extract noun-phrases as word bi-grams and tri-grams from the static document corpus and then models the topics using Latent Dirichlet Allocation algorithm. Additionally, we show that a small extension to our proposed framework can better rank documents in a large collection, which is a well studied area in information retrieval. Experiments conducted on three real world datasets show that this proposed framework outperforms state-of-the-art methods used for extracting concepts and ranking documents. When compared with the baselines chosen, our proposed concept extraction method showed an increased f-measure in the range of 16.65% to 22.04% and the proposed topic modeling guided document retrieval method showed 7.6%–16.61% increase in f-measure.
Cite this Research Publication : V. S. Anoop, S. Asharaf, P. Deepak, Topic Modeling for Unsupervised Concept Extraction and Document Ranking, Advances in Intelligent Systems and Computing, Springer International Publishing, 2017, https://doi.org/10.1007/978-3-319-68385-0_11