Improving Kurdish Web Mining through Tree Data Structure and Porter's Stemmer Algorithms

Ari M Saeed, Tarik A Rashid, Arazo M Mustafa, Polla Fattah, Birzo Ismael

January 01, 2016

Article

UKH Journal of Science and Engineering

1 min read

Abstract

Stemming is one of the main important preprocessing techniques that can be used to enhance the accuracy of text classification. The key purpose of using the stemming is combining the number of words that have same stem to decrease high dimensionality of feature space. Reducing feature space cause to decline time to construct a model and minimize the memory space. In this paper, a new stemming approach is explored for enhancing Kurdish text classification performance. Tree data structure and Porter’s stemmer algorithms are incorporated for building the proposed approach. The system is assessed through using Support Vector Machine (SVM) and Decision Tree (C4. 5) to illustrate the performance of the suggested stemmer after and before applying it. Furthermore, the usefulness of using stop words are considered before and after implementing the suggested approach.

Citation

Saeed, Ari M., Tarik A. Rashid, Arazo M. Mustafa, Polla Fattah, and Birzo Ismael. “Improving Kurdish web mining through tree data structure and Porter’s Stemmer algorithms.” UKH Journal of Science and Engineering 2, no. 1 (2018); 48-54.

Extra info

Type: Article
Link: UKH Journal of Science and Engineering

Polla Fattah