Neidio i’r brif dudalen lywio Neidio i chwilio Neidio i’r prif gynnwys

Threatening language detection from Urdu data with deep sequential model

  • Ashraf Ullah*
  • , Khair Ullah Khan
  • , Aurangzeb Khan
  • , Sheikh Tahir Bakhsh
  • , Atta Ur Rahman
  • , Sajida Akbar
  • , Bibi Saqia
  • , Toqir Rana (Golygydd)
  • *Awdur cyfatebol y gwaith hwn

Allbwn ymchwil: Cyfraniad at gyfnodolynErthygladolygiad gan gymheiriaid

8 Dyfyniadau (Scopus)

Crynodeb

The Urdu language is spoken and written on different social media platforms like Twitter, WhatsApp, Facebook, and YouTube. However, due to the lack of Urdu Language Processing (ULP) libraries, it is quite challenging to identify threats from textual and sequential data on the social media provided in Urdu. Therefore, it is required to preprocess the Urdu data as efficiently as English by creating different stemming and data cleaning libraries for Urdu data. Different lexical and machine learning-based techniques are introduced in the literature, but all of these are limited to the unavailability of online Urdu vocabulary. This research has introduced Urdu language vocabulary, including a stop words list and a stemming dictionary to preprocess Urdu data as efficiently as English. This reduced the input size of the Urdu language sentences and removed redundant and noisy information. Finally, a deep sequential model based on Long Short-Term Memory (LSTM) units is trained on the efficiently preprocessed, evaluated, and tested. Our proposed methodology resulted in good prediction performance, i.e., an accuracy of 82%, which is greater than the existing methods.
Iaith wreiddiolSaesneg
Rhif yr erthygle0290915
Tudalennau (o-i)e0290915
CyfnodolynPLoS ONE
Cyfrol19
Rhif cyhoeddi6
Dyddiad ar-lein cynnar6 Meh 2024
Dynodwyr Gwrthrych Digidol (DOIs)
StatwsCyhoeddwyd - 6 Meh 2024

Dyfynnu hyn