A Novel Framework for Text Preprocessing using NLP Approaches and Classification using Random Forest Grid Search Technique for Sentiment Analysis
Santosh Kumar et al.
What the paper says
Text preprocessing is a process to organize and prepare raw text for machine learning (ML) and deep learning (DL) models.This process is most widely used in Sentiment Analysis (SA).To process raw text data, Natural Language Processing (NLP) plays a vital role in data preprocessing by cleaning text, removing punctuations and stopwords, stemming and lemmatization.The classification of text data is a common use of NLP.An effective data preprocessing approach also identifies text features efficiently.The most essential component of text classification is feature extraction from raw text data to input into ML and DL models.Feature extraction prepares training data sets effectively for ML and DL models to find good results.In this paper, we proposed a novel framework for text preprocessing using NLP approaches, and feature extraction, and expanded it for the Random Forest ML classifier.To improve the performance of the random forest classifier, we examined the grid search technique with estimator Random Forest Classifier and measured the performance of the model.The accuracy measure for the random forest model was 93%, while the accuracy measure using the grid search technique was 94%, which shows that the grid search technique enhances the model performance.
2 citations
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.25 × 0.4 = 0.10 |
| M · momentum | 0.55 × 0.15 = 0.08 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.