Simple Optimal Sampling Algorithm to Strengthen Digital Soil Mapping Using the Spatial Distribution of Machine Learning Predictive Uncertainty: A Case Study for Field Capacity Prediction

被引:6
作者
Yang, Hyunje [1 ]
Lim, Honggeun [1 ]
Moon, Haewon [1 ]
Li, Qiwen [1 ]
Nam, Sooyoun [1 ]
Kim, Jaehoon [1 ]
Choi, Hyung Tae [1 ]
机构
[1] Natl Inst Forest Sci, Forest Environm & Conservat Dept, Seoul 02455, South Korea
关键词
digital soil mapping; field capacity; machine learning; predictive uncertainty; sample site survey; soil investigation plan; K-NEAREST NEIGHBOR; ORGANIC-MATTER; TREE; CLASSIFICATION; CARBON;
D O I
10.3390/land11112098
中图分类号
X [环境科学、安全科学];
学科分类号
08 ; 0830 ;
摘要
Machine learning models are now capable of delivering coveted digital soil mapping (DSM) benefits (e.g., field capacity (FC) prediction); therefore, determining the optimal sample sites and sample size is essential to maximize the training efficacy. We solve this with a novel optimal sampling algorithm that allows the authentic augmentation of insufficient soil features using machine learning predictive uncertainty. Nine hundred and fifty-three forest soil samples and geographically referenced forest information were used to develop predictive models, and FCs in South Korea were estimated with six predictor set hierarchies. Random forest and gradient boosting models were used for estimation since tree-based models had better predictive performance than other machine learning algorithms. There was a significant relationship between model predictive uncertainties and training data distribution, where higher uncertainties were distributed in the data scarcity area. Further, we confirmed that the predictive uncertainties decreased when additional sample sites were added to the training data. Environmental covariate information of each grid cell in South Korea was then used to select the sampling sites. Optimal sites were coordinated at the cell having the highest predictive uncertainty, and the sample size was determined using the predictable rate. This intuitive method can be generalized to improve global DSM.
引用
收藏
页数:18
相关论文
共 43 条
[21]   Classification of remotely sensed imagery using stochastic gradient boosting as a refinement of classification tree analysis [J].
Lawrence, R ;
Bunn, A ;
Powell, S ;
Zambon, M .
REMOTE SENSING OF ENVIRONMENT, 2004, 90 (03) :331-336
[22]   Evaluation of Rainfall Erosivity Factor Estimation Using Machine and Deep Learning Models [J].
Lee, Jimin ;
Lee, Seoro ;
Hong, Jiyeong ;
Lee, Dongjun ;
Bae, Joo Hyun ;
Yang, Jae E. ;
Kim, Jonggun ;
Lim, Kyoung Jae .
WATER, 2021, 13 (03)
[23]   Development of Pedo-Transfer Functions for the Saturated Hydraulic Conductivity of Forest Soil in South Korea Considering Forest Stand and Site Characteristics [J].
Lim, Honggeun ;
Yang, Hyunje ;
Chun, Kun Woo ;
Choi, Hyung Tae .
WATER, 2020, 12 (08)
[24]   Digital mapping of soil properties in Canadian managed forests at 250 m of resolution using the k-nearest neighbor method [J].
Mansuy, Nicolas ;
Thiffault, Evelyne ;
Pare, David ;
Bernier, Pierre ;
Guindon, Luc ;
Villemaire, Philippe ;
Poirier, Vincent ;
Beaudoin, Andre .
GEODERMA, 2014, 235 :59-73
[25]   A QUANTITATIVE AUSTRALIAN APPROACH TO MEDIUM AND SMALL-SCALE SURVEYS BASED ON SOIL STRATIGRAPHY AND ENVIRONMENTAL CORRELATION [J].
MCKENZIE, NJ ;
AUSTIN, MP .
GEODERMA, 1993, 57 (04) :329-355
[26]   Inverse method using boosted regression tree and k-nearest neighbor to quantify effects of point and non-point source nitrate pollution in groundwater [J].
Motevalli, Alireza ;
Naghibi, Seyed Amir ;
Hashemi, Hossein ;
Berndtsson, Ronny ;
Pradhan, Biswajeet ;
Gholami, Vahid .
JOURNAL OF CLEANER PRODUCTION, 2019, 228 :1248-1263
[27]   An introduction to decision tree modeling [J].
Myles, AJ ;
Feudale, RN ;
Liu, Y ;
Woody, NA ;
Brown, SD .
JOURNAL OF CHEMOMETRICS, 2004, 18 (06) :275-285
[28]   Efficient partition of integer optimization problems with one-hot encoding [J].
Okada, Shuntaro ;
Ohzeki, Masayuki ;
Taguchi, Shinichiro .
SCIENTIFIC REPORTS, 2019, 9 (1)
[29]   Organic carbon, organic matter and bulk density relationships in boreal forest soils [J].
Perie, Catherine ;
Ouimet, Rock .
CANADIAN JOURNAL OF SOIL SCIENCE, 2008, 88 (03) :315-325
[30]   Perspectives on validation in digital soil mapping of continuous attributes-A review [J].
Piikki, Kristin ;
Wetterlind, Johanna ;
Soderstrom, Mats ;
Stenberg, Bo .
SOIL USE AND MANAGEMENT, 2021, 37 (01) :7-21