A geometric framework for outlier detection in high-dimensional data

被引:0
作者
Herrmann, Moritz [1 ]
Pfisterer, Florian [1 ]
Scheipl, Fabian [1 ]
机构
[1] Ludwig Maximilians Univ Munchen, Dept Stat, Ludwigstr 33, D-80539 Munich, Germany
关键词
anomaly detection; dimension reduction; manifold learning; outlier detection; REDUCTION;
D O I
10.1002/widm.1491
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Outlier or anomaly detection is an important task in data analysis. We discuss the problem from a geometrical perspective and provide a framework which exploits the metric structure of a data set. Our approach rests on the manifold assumption, that is, that the observed, nominally high-dimensional data lie on a much lower dimensional manifold and that this intrinsic structure can be inferred with manifold learning methods. We show that exploiting this structure significantly improves the detection of outlying observations in high dimensional data. We also suggest a novel, mathematically precise and widely applicable distinction between distributional and structural outliers based on the geometry and topology of the data manifold that clarifies conceptual ambiguities prevalent throughout the literature. Our experiments focus on functional data as one class of structured high-dimensional data, but the framework we propose is completely general and we include image and graph data applications. Our results show that the outlier structure of highdimensional and non-tabular data can be detected and visualized using manifold learning methods and quantified using standard outlier scoring methods applied to the manifold embedding vectors. This article is categorized under: Technologies > Structure Discovery and Clustering Fundamental Concepts of Data and Knowledge > Data Concepts Technologies > Visualization
引用
收藏
页数:20
相关论文
共 50 条
  • [31] A survey of outlier detection in high dimensional data streams
    Souiden, Imen
    Omri, Mohamed Nazih
    Brahmi, Zaki
    COMPUTER SCIENCE REVIEW, 2022, 44
  • [32] A NOVEL TENSOR ALGEBRAIC APPROACH FOR HIGH-DIMENSIONAL OUTLIER DETECTION UNDER DATA MISALIGNMENT
    Fan, Bo
    Zhang, Zemin
    Aeron, Shuchin
    2016 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP), 2016, : 3628 - 3632
  • [33] PCA leverage: outlier detection for high-dimensional functional magnetic resonance imaging data
    Mejia, Amanda F.
    Nebel, Mary Beth
    Eloyan, Ani
    Caffo, Brian
    Lindquist, Martin A.
    BIOSTATISTICS, 2017, 18 (03) : 521 - 536
  • [34] Support high-order tensor data description for outlier detection in high-dimensional big sensor data
    Deng, Xiaowu
    Jiang, Peng
    Peng, Xiaoning
    Mi, Chunqiao
    FUTURE GENERATION COMPUTER SYSTEMS-THE INTERNATIONAL JOURNAL OF ESCIENCE, 2018, 81 : 177 - 187
  • [35] Outlier Detection based on Sparse Coding and Neighbor Entropy in High-dimensional Space
    Gu, Ping
    Chow, Meng
    Shao, Siyu
    17TH ACM INTERNATIONAL CONFERENCE ON COMPUTING FRONTIERS 2020 (CF 2020), 2020, : 202 - 207
  • [36] A High-dimensional Outlier Detection Algorithm Base on Relevant Subspace
    Gao, Zhipeng
    Zhao, Yang
    Niu, Kun
    Fan, Yidan
    2017 IEEE 15TH INTL CONF ON DEPENDABLE, AUTONOMIC AND SECURE COMPUTING, 15TH INTL CONF ON PERVASIVE INTELLIGENCE AND COMPUTING, 3RD INTL CONF ON BIG DATA INTELLIGENCE AND COMPUTING AND CYBER SCIENCE AND TECHNOLOGY CONGRESS(DASC/PICOM/DATACOM/CYBERSCI, 2017, : 1001 - 1008
  • [37] VOA*: Fast Angle-Based Outlier Detection over High-Dimensional Data Streams
    Khalique, Vijdan
    Kitagawa, Hiroyuki
    ADVANCES IN KNOWLEDGE DISCOVERY AND DATA MINING, PAKDD 2021, PT I, 2021, 12712 : 40 - 52
  • [38] Anomaly Detection in High-Dimensional Data
    Talagala, Priyanga Dilini
    Hyndman, Rob J.
    Smith-Miles, Kate
    JOURNAL OF COMPUTATIONAL AND GRAPHICAL STATISTICS, 2021, 30 (02) : 360 - 374
  • [39] A High-Dimensional Outlier Detection Approach Based on Local Coulomb Force
    Zhu, Pengyun
    Zhang, Chaowei
    Li, Xiaofeng
    Zhang, Jifu
    Qin, Xiao
    IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2023, 35 (06) : 5506 - 5520
  • [40] An Ensemble Outlier Detection Method Based on Information Entropy-Weighted Subspaces for High-Dimensional Data
    Li, Zihao
    Zhang, Liumei
    ENTROPY, 2023, 25 (08)