Machine Learning and Feature Selection for soil spectroscopy. An evaluation of Random Forest wrappers to predict soil organic matter, clay, and carbonates

被引:4
|
作者
Canero, Francisco M. [1 ]
Rodriguez-Galiano, Victor [1 ]
Aragones, David [2 ]
机构
[1] Univ Seville, Dept Phys Geog & Reg Geog Anal, Seville 41004, Spain
[2] CSIC, Remote Sensing & Geog Informat Syst Lab LAST EBD, Donana Biol Stn, Seville 41092, Spain
关键词
Random forest; Sequential flotant selection; Sequential flotant forward selection; Partial least squares regression; Wrapper methods; Sierra de las nieves; PARTIAL LEAST-SQUARES; DIFFUSE-REFLECTANCE SPECTROSCOPY; INFRARED SPECTROSCOPY; HYPERSPECTRAL IMAGES; PRINCIPAL COMPONENT; TOTAL NITROGEN; REGRESSION; AIRBORNE; CLASSIFICATION; PLSR;
D O I
10.1016/j.heliyon.2024.e30228
中图分类号
O [数理科学和化学]; P [天文学、地球科学]; Q [生物科学]; N [自然科学总论];
学科分类号
07 ; 0710 ; 09 ;
摘要
Soil spectroscopy estimates soil properties using the absorption features in soil spectra. However, modelling soil properties with soil spectroscopy is challenging due to the high dimensionality of spectral data. Feature Selection wrapper methods are promising approaches to reduce the dimensionality but are barely used in soil spectroscopy. The aim of this study is to evaluate the performance of two feature selection wrapper methods, Sequential Forward Selection (SFS) and Sequential Flotant Forward Selection (SFFS) built using the Random Forest (RF) algorithm, for dimensionality reduction of spectral data and predictive modelling of modelling soil organic matter (SOM), clay and carbonates. The reflectance of 100 soil samples, acquired from Sierra de las Nieves (Spain), was measured under laboratory conditions using ASD FieldSpec Pro JR. Four different datasets were obtained after applying two spectral preprocessing methods to raw spectra: raw spectra, Continuum Removal (CR), Multiplicative Scatter Correction (MSC), and a socalled " Global " dataset composed of raw, CR and MSC features. The performance of RF models built with feature selection methods was compared to that of Partial Least Squares Regression (PLSR) and RF (alone). RF models built with SFS and SFFS outperformed PLSR and RF alone models: The best RF models with feature selection had a respective ratio of performance to interquartile distance of 1.93, 0.38 and 2.56. PLSR models had an accuracy of 1.41, 0.29 and 1.81 for SOM, carbonates, and clay, respectively. RF alone had a respective performance of 1.29, 0.29 and 1.81. The application of feature selection wrapper methods reduced the number of features to less than 1 % of the starting features. Features were selected across all spectra for SOM and clay, and around 900 nm, 1900 nm, and 2350 nm for carbonates. However, feature selection highlighted features around 1100 nm in SOM modelling, as well as other features around 2200 nm, which is considered a main absorption feature of clay. The application of feature selection with Random Forest was very important in improving modelling accuracy, reducing the redundant features and avoiding the curse of dimensionality or Hughes effect. Thus, this research showed an alternative to dimensionality reduction approaches that have been applied to date to model soil properties with spectroscopy and paves the way for further scientific investigation based on feature selection methods and machine learning.
引用
收藏
页数:19
相关论文
共 50 条
  • [1] Evaluation of Machine Learning Approaches to Predict Soil Organic Matter and pH Using vis-NIR Spectra
    Yang, Meihua
    Xu, Dongyun
    Chen, Songchao
    Li, Hongyi
    Shi, Zhou
    SENSORS, 2019, 19 (02):
  • [2] Artificial bee colony feature selection algorithm combined with machine learning algorithms to predict vertical and lateral distribution of soil organic matter in South Dakota, USA
    Taghizadeh-Mehrjardi, Ruhollah
    Neupane, Ram
    Sood, Kunal
    Kumar, Sandeep
    CARBON MANAGEMENT, 2017, 8 (03) : 277 - 291
  • [3] Combination of machine learning and VIRS for predicting soil organic matter
    Dong, Zhenyu
    Wang, Ni
    Liu, Jinbao
    Xie, Jiancang
    Han, Jichang
    JOURNAL OF SOILS AND SEDIMENTS, 2021, 21 (07) : 2578 - 2588
  • [4] Predicting nickel concentration in soil using reflectance spectroscopy associated with organic matter and clay minerals
    Sun, Weichao
    Zhang, Xia
    Sun, Xuejian
    Sun, Yanli
    Cen, Yi
    GEODERMA, 2018, 327 : 25 - 35
  • [5] Soil organic matter and clay predictions by laboratory spectroscopy: Data spatial correlation
    da Silva-Sangoi, Daniely Vaz
    Horst, Taciara Zborowski
    Moura-Bueno, Jean Michel
    Diniz Dalmolin, Ricardo Simao
    Sebem, Elodio
    Gebler, Luciano
    Santos, Marcio da Silva
    GEODERMA REGIONAL, 2022, 28
  • [6] Soil salinization prediction through feature selection and machine learning at the irrigation district scale
    Xie, Junbo
    Shi, Cong
    Liu, Yang
    Wang, Qi
    Zhong, Zhibo
    He, Shuai
    Wang, Xingpeng
    FRONTIERS IN EARTH SCIENCE, 2025, 12
  • [7] Random forest prediction model for the soil organic matter with optimized spectral inputs
    Zhang X.
    Meng X.
    Tang H.
    Liu H.
    Zhang X.
    Liu Q.
    Nongye Gongcheng Xuebao/Transactions of the Chinese Society of Agricultural Engineering, 2023, 39 (02): : 90 - 99
  • [8] Evaluating Airborne Hyperspectral Scanner (AHS) for the mapping of soil organic matter and clay in a Mediterranean forest ecosystem
    Canero, Francisco M.
    Rodriguez-Galiano, Victor
    Chabrillat, Sabine
    CATENA, 2025, 252
  • [9] Estimating Soil Organic Matter Content Using Sentinel-2 Imagery by Machine Learning in Shanghai
    Wang, Xinxin
    Han, Jigang
    Wang, Xia
    Yao, Huaiying
    Zhang, Lang
    IEEE ACCESS, 2021, 9 : 78215 - 78225
  • [10] On-line vis-NIR spectroscopy prediction of soil organic carbon using machine learning
    Nawar, S.
    Mouazen, A. M.
    SOIL & TILLAGE RESEARCH, 2019, 190 : 120 - 127