WOLI at SemEval-2020 Task 12: Arabic Offensive Language Identification on Different Twitter Datasets

Otiefy, Yasser; Abdelmalek, Ahmed; Hosary, Islam El

Computer Science > Computation and Language

arXiv:2009.05456 (cs)

[Submitted on 11 Sep 2020]

Title:WOLI at SemEval-2020 Task 12: Arabic Offensive Language Identification on Different Twitter Datasets

Authors:Yasser Otiefy (WideBot), Ahmed Abdelmalek (WideBot), Islam El Hosary (WideBot)

View PDF

Abstract:Communicating through social platforms has become one of the principal means of personal communications and interactions. Unfortunately, healthy communication is often interfered by offensive language that can have damaging effects on the users. A key to fight offensive language on social media is the existence of an automatic offensive language detection system. This paper presents the results and the main findings of SemEval-2020, Task 12 OffensEval Sub-task A Zampieri et al. (2020), on Identifying and categorising Offensive Language in Social Media. The task was based on the Arabic OffensEval dataset Mubarak et al. (2020). In this paper, we describe the system submitted by WideBot AI Lab for the shared task which ranked 10th out of 52 participants with Macro-F1 86.9% on the golden dataset under CodaLab username "yasserotiefy". We experimented with various models and the best model is a linear SVM in which we use a combination of both character and word n-grams. We also introduced a neural network approach that enhanced the predictive ability of our system that includes CNN, highway network, Bi-LSTM, and attention layers.

Subjects:	Computation and Language (cs.CL); Social and Information Networks (cs.SI)
Cite as:	arXiv:2009.05456 [cs.CL]
	(or arXiv:2009.05456v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2009.05456

Submission history

From: Yasser Otiefy [view email]
[v1] Fri, 11 Sep 2020 14:10:03 UTC (81 KB)

Computer Science > Computation and Language

Title:WOLI at SemEval-2020 Task 12: Arabic Offensive Language Identification on Different Twitter Datasets

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:WOLI at SemEval-2020 Task 12: Arabic Offensive Language Identification on Different Twitter Datasets

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators