A Cooperative Binary-Clustering Framework Based on Majority Voting for Twitter Sentiment Analysis

Maryum Bibi, Wajid Aziz*, Majid Almaraashi, Imtiaz Hussain Khan, Malik Sajjad Ahmed Nadeem, Nazneen Habib

*Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

39 Citations (Scopus)

Abstract

Twitter sentiment analysis is a challenging problem in natural language processing. For this purpose, supervised learning techniques have mostly been employed, which require labeled data for training. However, it is very time consuming to label datasets of large size. To address this issue, unsupervised learning techniques such as clustering can be used. In this study, we explore the possibility of using hierarchical clustering for twitter sentiment analysis. Three hierarchical-clustering techniques, namely single linkage (SL), complete linkage (CL) and average linkage (AL), are examined. A cooperative framework of SL, CL and AL is built to select the optimal cluster for tweets wherein the notion of optimal-cluster selection is operationalized using majority voting. The hierarchical clustering techniques are also compared with k-means and two state-of-the-art classifiers (SVM and Naïve Bayes). The performance of clustering and classification is measured in terms of accuracy and time efficiency. The experimental results indicate that cooperative clustering based on majority voting approach is robust in terms of good quality clusters with tradeoff of poor time efficiency. The results also suggest that the accuracy of the proposed clustering framework is comparable to classifiers which is encouraging.

Original languageEnglish
Article number9049112
Pages (from-to)68580-68592
Number of pages13
JournalIEEE Access
Volume8
DOIs
Publication statusPublished - 27 Mar 2020
Externally publishedYes

Keywords

  • Cooperative clustering
  • majority voting
  • sentiment analysis
  • twitter sentiment analysis

Cite this