Resource Time Series Analysis and Forecasting in Large-Scale Virtual Clusters

被引:0
作者
Lin, Yue [1 ,2 ]
Wen, Jiamin [3 ]
Zhang, Xudong [4 ]
Liang, Yan [3 ]
Li, Jianjiang [1 ]
机构
[1] Univ Sci & Technol Beijing, Dept Comp Sci & Technol, Beijing 100083, Peoples R China
[2] 41st Inst CETC, Qingdao 266555, Peoples R China
[3] China Natl Petr Corp, BGP Inc, Zhuozhou 072751, Peoples R China
[4] Natl Engn Res Ctr Oil & Gas Explorat Comp Software, Zhuozhou 072751, Peoples R China
来源
BIG DATA MINING AND ANALYTICS | 2025年 / 8卷 / 03期
关键词
workload forecasting; multivariate time series forecasting; deep learning; MODEL; PREDICTION;
D O I
10.26599/BDMA.2024.9020085
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
In today's rapidly evolving internet landscape, prominent companies across various industries face increasingly complex business operations, leading to significant cluster-scale growth. However, this growth brings about challenges in cluster management and the inefficient utilization of vast amounts of data due to its low value density. This paper, based on the large-scale cluster virtualization and monitoring system of the data center of the Bureau of Geophysical Prospecting (BGP), utilizes time series data of host resources from the monitoring system's time series database to propose a multivariate multi-step time series forecasting model, MUL-CNN-BiGRU-Attention, for forecasting CPU load on virtual cluster hosts. The model undergoes extensive offline training using a large volume of time series data, followed by deployment using TensorFlow Serving. Recent small-batch data are employed for fine-tuning model parameters to better adapt to current data patterns. Comparative experiments are conducted between the proposed model and other baseline models, demonstrating notable improvements in Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and $R<^>{2}$ metrics by up to 35.2%, 56.1%, 32.5%, and 10.3%, respectively. Additionally, ablation experiments are designed to investigate the impact of different factors on the performance of the forecasting model, providing valuable insights for parameter optimization based on experimental results.
引用
收藏
页码:592 / 605
页数:14
相关论文
共 32 条
  • [31] [杨海民 Yang Haimin], 2019, [计算机科学, Computer Science], V46, P21
  • [32] Handling missing data in near real-time environmental monitoring: A system and a review of selected methods
    Zhang, Yifan
    Thorburn, Peter J.
    [J]. FUTURE GENERATION COMPUTER SYSTEMS-THE INTERNATIONAL JOURNAL OF ESCIENCE, 2022, 128 : 63 - 72