☆ 4.6 Article

A survey on Urdu and Urdu like language stemmers and stemming techniques

ARTIFICIAL INTELLIGENCE REVIEW (2018)

Journal

ARTIFICIAL INTELLIGENCE REVIEW

Volume 49, Issue 3, Pages 339-373

Publisher

SPRINGER

DOI: 10.1007/s10462-016-9527-1

Keywords

Stemming; Natural Language Processing; Information Retrieval; Urdu; Suffixes; Stemming Techniques

Ask authors/readers for more resources

Protocol

Community support

Reagent

Community support

Abstract

Stemming is one of the basic steps in natural language processing applications such as information retrieval, parts of speech tagging, syntactic parsing and machine translation, etc. It is a morphological process that intends to convert the inflected forms of a word into its root form. Urdu is a morphologically rich language, emerged from different languages, that includes prefix, suffix, infix, co-suffix and circumfixes in inflected and multi-gram words that need to be edited in order to convert them into their stems. This editing (insertion, deletion and substitution) makes the stemming process difficult due to language morphological richness and inclusion of words of foreign languages like Persian and Arabic. In this paper, we present a comprehensive review of different algorithms and techniques of stemming Urdu text and also considering the syntax, morphological similarity and other common features and stemming approaches used in Urdu like languages, i.e. Arabic and Persian analyzed, extract main features, merits and shortcomings of the used stemming approaches. In this paper, we also discuss stemming errors, basic difference between stemming and lemmatization and coin a metric for classification of stemming algorithms. In the final phase, we have presented the future work directions.

A survey on Urdu and Urdu like language stemmers and stemming techniques

Journal

ARTIFICIAL INTELLIGENCE REVIEW

Publisher

SPRINGER

Keywords

Categories

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

A survey on Urdu and Urdu like language stemmers and stemming techniques

Journal

ARTIFICIAL INTELLIGENCE REVIEW

Publisher

SPRINGER

Keywords

Categories

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Export Citation

Share Paper