咨询与建议

看过本文的还看了

相关文献

该作者的其他文献

文献详情 >Topic modeling Twitter data us... 收藏

Topic modeling Twitter data using Latent Dirichlet Allocation and Latent Semantic Analysis

作     者:Siti Qomariyah Nur Iriawan Kartika Fithriasari 

作者机构:Department of Statistics Faculty of Mathematics Computing and Data Science Institut Teknologi Sepuluh Nopember Kampus ITS-Sukolilo 60111 Surabaya Indonesia 

出 版 物:《AIP Conference Proceedings》 

年 卷 期:2019年第2194卷第1期

学科分类:07[理学] 0702[理学-物理学] 

摘      要:The industrial world has entered the era of industrial revolution 4.0. In this era, there is an urgent data requirement from the community to support service policies. Because of that, Surabaya Government made Media Center Surabaya. This media is used to accommodate all the aspiration of Surabaya citizen. To access this media, a citizen can use Twitter. The topic which is discussed in Twitter is important information that we need to know. The information can be used to improve the performance of Surabaya Government services. Twitter data is a text data that consists of thousands of variables. Text mining is frequently used to analyze this kind of data, including topic modeling and sentiment analysis. This study would work on topic modeling focused on the algorithm employing Latent Dirichlet Allocation (LDA) and Latent Semantic Analysis (LSA). The evaluation of the algorithm performance uses the topic coherence. As unstructured data, the Twitter data need preprocessing before the analysis. The stages of preprocessing include cleansing, stemming, and stop words. The advantages of LSA are fast and easy to implement. LSA, on the other hand, doesn’t consider the relationship between documents in the corpus, while LDA does. This study shows that LDA gives a better result than LSA.

读者评论 与其他读者分享你的观点

用户名:未登录
我的评分