当前位置: X-MOL 学术Mach. Learn. › 论文详情
Our official English website, www.x-mol.net, welcomes your feedback! (Note: you will need to create a separate account there.)
Embed2Detect: temporally clustered embedded words for event detection in social media
Machine Learning ( IF 4.3 ) Pub Date : 2021-05-24 , DOI: 10.1007/s10994-021-05988-7
Hansi Hettiarachchi , Mariam Adedoyin-Olowe , Jagdev Bhogal , Mohamed Medhat Gaber

Social media is becoming a primary medium to discuss what is happening around the world. Therefore, the data generated by social media platforms contain rich information which describes the ongoing events. Further, the timeliness associated with these data is capable of facilitating immediate insights. However, considering the dynamic nature and high volume of data production in social media data streams, it is impractical to filter the events manually and therefore, automated event detection mechanisms are invaluable to the community. Apart from a few notable exceptions, most previous research on automated event detection have focused only on statistical and syntactical features in data and lacked the involvement of underlying semantics which are important for effective information retrieval from text since they represent the connections between words and their meanings. In this paper, we propose a novel method termed Embed2Detect for event detection in social media by combining the characteristics in word embeddings and hierarchical agglomerative clustering. The adoption of word embeddings gives Embed2Detect the capability to incorporate powerful semantical features into event detection and overcome a major limitation inherent in previous approaches. We experimented our method on two recent real social media data sets which represent the sports and political domain and also compared the results to several state-of-the-art methods. The obtained results show that Embed2Detect is capable of effective and efficient event detection and it outperforms the recent event detection methods. For the sports data set, Embed2Detect achieved 27% higher F-measure than the best-performed baseline and for the political data set, it was an increase of 29%.



中文翻译:

Embed2Detect:用于在社交媒体中进行事件检测的时间聚集的嵌入式单词

社交媒体正成为讨论世界各地正在发生的事情的主要媒介。因此,社交媒体平台生成的数据包含描述正在进行的事件的丰富信息。此外,与这些数据相关的及时性能够促进即时见解。但是,考虑到社交媒体数据流中数据的动态性质和大量数据,手动过滤事件是不切实际的,因此,自动事件检测机制对社区而言是无价的。除了一些值得注意的例外,以前有关自动事件检测的大多数研究仅集中于数据中的统计和句法特征,并且缺少底层语义的介入,这对于从文本有效检索信息非常重要,因为它们代表了单词及其含义之间的联系。在本文中,我们提出了一种新的方法,称为Embed2Detect通过结合词嵌入和层次化聚集聚类的特征,在社交媒体中进行事件检测。单词嵌入的采用使Embed2Detect能够将强大的语义特征整合到事件检测中,并克服了先前方法固有的主要限制。我们在代表体育和政治领域的两个最新的真实社交媒体数据集上实验了我们的方法,并将结果与​​几种最新方法进行了比较。获得的结果表明Embed2Detect能够进行有效且高效的事件检测,并且胜过最近的事件检测方法。对于体育数据集,Embed2Detect的F值比最佳执行基准高27%;对于政治数据集,F-measure值提高了29%。

更新日期:2021-05-25
down
wechat
bug