Knowledge expansion of metadata using script mining analysis in multimedia recommendation

被引：0

作者：

Joo-Chang Kim

Kyung-Yong Chung

机构：

[1] Kyonggi University,Data Mining Lab., Department of Computer Science

[2] Kyonggi University,Division of Computer Science and Engineering

来源：

Multimedia Tools and Applications | 2021年 / 80卷

关键词：

Multimedia; Data mining; Knowledge discovery; Recommendation; Metadata;

D O I：

暂无

中图分类号：

学科分类号：

摘要：

In this paper, a method for knowledge expansion of metadata using script mining analysis for multimedia recommendation systems is proposed. The method allows the extraction of new metadata and knowledge expansion through the mining analysis of multimedia scripts, which include a large amount of information. The scripts are collected by a Web crawler based on Python. From the collected scripts, hidden information is extracted through keyword analysis and sentiment analysis. In keyword analysis, scripts, unlike general documents, show a high frequency of names of characters or proper nouns. Such names or proper nouns are not frequently used in other media content, and therefore, their importance is high. Frequently, they are already offered in the conventional metadata, and consequently cause information duplication. Accordingly, term frequency–inverse document and metadata frequency (TF–IDMF), which considers the frequency of metadata in general term frequency–inverse document frequency (TF–IDF), is used. Thus, the importance of the names of characters or proper nouns in scripts can be decreased. Because the keywords for the extracted scripts are in fact included in the scripts, they can be used for precise multimedia search and recommendation. In sentiment analysis, the AFINN lexicon and the Bing lexicon are utilized to scan words in a script. The Bing lexicon is used to examine whether the words in the entire script are positive or negative. Then, the total numbers of positive words and negative words are used to calculate the representative sentiment of the script. The AFINN lexicon includes approximately 170 sentiment words, the negative or positive sentiment of which is presented in the range − 5 to +5. One script is divided into 100 sentences, and then, the representative sentiment in each sentence is evaluated as either positive or negative. Through script scanning, the flow of sentiment in multimedia streams can be discovered. The Bing lexicon categorizes words into positive, negative, and neutral sentiments. Through script scanning, the words included in each category can be quantified. Depending on the result of the script sentiment analysis, a different sentence embedding method based on inter-sentence similarity is used to cluster similar media. The results of the keyword analysis and sentiment analysis of a script are added to the metadata in a new column in a knowledge base to expand knowledge. To evaluate the significance of multimedia recommendations, keywords and sentiment information are used, and then, the similarity and clustering of the extracted media are assessed. As a result, script mining analysis based on the attributes that include actual information of media is considerably better than that based on types or a range of metadata attributes. Therefore, the proposed knowledge expansion method achieves significant results and shows an excellent performance in multimedia recommendation.

引用

页码：34679 / 34695

页数：16

共 63 条

[1] Aghdam MH(2019)Context-aware recommender systems using hierarchical hidden Markov model Physica A: Statistical Mechanics and its Applications 518 89-98
[2] Baek JW(2019)Hybrid clustering based health decision-making for improving dietary habits Technol Health Care 27 459-472
[3] Kim JC(2013)Recommender systems survey Knowl-Based Syst 46 109-132
[4] Chun J(2016)Adaptive subframe allocation for next generation multimedia delivery over hybrid LTE unicast broadcast IEEE Trans Broadcast 62 540-551
[5] Chung K(2010)Metadata models of the world wide web Libr Technol Rep 46 12-19
[6] Bobadilla J(2009)A survey of text mining techniques and applications J Emerg Technol Web Intell 1 60-76
[7] Ortega F(2016)Knowledge-based dietary nutrition recommendation for obese management Inf Technol Manag 17 29-42
[8] Hernando A(2002)An efficient k-means clustering algorithm: analysis and implementation IEEE Trans Pattern Anal Mach Intell 7 881-892
[9] Gutiérrez A(2017)Emerging risk forecast system using associative index mining analysis Clust Comput 20 547-558
[10] Christodoulou L(2018)Mining health-risk factors using PHR similarity in a hybrid P2P network Peer-to-Peer Netw Appl 11 1278-1287

← 1 2 3 4 5 6 7 →