Fine-Grained Representation Learning and Recognition by Exploiting Hierarchical Semantic Embedding

被引:75
作者
Chen, Tianshui [1 ]
Wu, Wenxi [1 ]
Gao, Yuefang [2 ]
Dong, Le [3 ]
Luo, Xiaonan [4 ]
Lin, Liang [1 ]
机构
[1] Sun Yat Sen Univ, Guangzhou, Guangdong, Peoples R China
[2] South China Agr Univ, Guangzhou, Guangdong, Peoples R China
[3] Univ Elect Sci & Technol China, Chengdu, Sichuan, Peoples R China
[4] Guilin Univ Elect Technol, Guilin, Peoples R China
来源
PROCEEDINGS OF THE 2018 ACM MULTIMEDIA CONFERENCE (MM'18) | 2018年
基金
美国国家科学基金会;
关键词
Semantic Embedding; Fine-Grained Image Recognition; Category Hierarchy; CNNS;
D O I
10.1145/3240508.3240523
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
Object categories inherently form a hierarchy with different levels of concept abstraction, especially for fine-grained categories. For example, birds (Aves) can be categorized according to a four-level hierarchy of order, family, genus, and species. This hierarchy encodes rich correlations among various categories across different levels, which can effectively regularize the semantic space and thus make prediction less ambiguous. However, previous studies of fine-grained image recognition primarily focus on categories of one certain level and usually overlook this correlation information. In this work, we investigate simultaneously predicting categories of different levels in the hierarchy and integrating this structured correlation information into the deep neural network by developing a novel Hierarchical Semantic Embedding (HSE) framework. Specifically, the HSE framework sequentially predicts the category score vector of each level in the hierarchy, from highest to lowest. At each level, it incorporates the predicted score vector of the higher level as prior knowledge to learn finer-grained feature representation. During training, the predicted score vector of the higher level is also employed to regularize label prediction by using it as soft targets of corresponding sub-categories. To evaluate the proposed framework, we organize the 200 bird species of the Caltech-UCSD birds dataset with the four-level category hierarchy and construct a large-scale butterfly dataset that also covers four level categories. Extensive experiments on these two and the newly-released VegFru datasets demonstrate the superiority of our HSE framework over the baseline methods and existing competitors.
引用
收藏
页码:2023 / 2031
页数:9
相关论文
共 50 条
[1]  
[Anonymous], 2015, P IEEE C COMP VIS PA
[2]  
[Anonymous], 2013, Tech. rep.
[3]  
[Anonymous], 2016, ARXIV160306765
[4]  
[Anonymous], 2017, IEEE C COMP VIS PATT
[5]  
[Anonymous], 2011, Technical Report CNS-TR-2011-001
[6]  
[Anonymous], 2011, P CVPR WORKSHOP FINE
[7]  
[Anonymous], 2018, P INT JOINT C ART IN
[8]  
[Anonymous], ARXIV161105109
[9]  
[Anonymous], P IEEE C COMP VIS PA
[10]  
[Anonymous], 2014, BRIT MACH VIS C