Improving Kurdish Web Mining through Tree Data Structure and Porter’ Stemmer Algorithms

Ari M. Saeed, Tarik A. Rashid, Arazo M. Mustafa, Polla Fattah, Birzo Ismaels

Abstract

Stemming is one of the main important preprocessing techniques that can be used to enhance the accuracy of text classification. The key purpose of using the stemming is combining the number of words that have same stem to decrease high dimensionality of feature space. Reducing feature space cause to decline time to construct a model and minimize the memory space. In this paper, a new stemming approach is explored for enhancing Kurdish text classification performance. Tree data structure and Porter’s stemmer algorithms are incorporated for building the proposed approach. The system is assessed through using Support Vector Machine (SVM) and Decision Tree (C4.5) to illustrate the performance of the suggested stemmer after and before applying it. Furthermore, the usefulness of using stop words are considered before and after implementing the suggested approach.

Citation

Extra info


Polla Fattah