نشریه سنجش از دور و GIS  ایران

نشریه سنجش از دور و GIS ایران

ارزیابی مقایسه‌ای الگوریتم‌های یادگیری ماشین با تأکید بر درختان فوق تصادفی (Extra Trees) و تفسیرپذیری SHAP در برآورد فرونشست زمین دشت‌های ایران

نوع مقاله : مقاله پژوهشی

نویسندگان
1 رئیس پژوهشکده سنجش از دور و GIS محیطی دانشگاه علوم کشاورزی و منابع طبیعی ساری
2 دانش آموخته دکتری، علوم و مهندسی آبخیزداری، دانشکده منابع طبیعی، دانشگاه منابع طبیعی ساری، ساری، ایران.
3 استادیار، گروه مرتع و آبخیزداری، دانشکده منابع طبیعی و محیط‌زیست، دانشگاه بیرجند، بیرجند، ایران.
4 دانش آموخته دکتری، علوم و مهندسی آبخیزداری، دانشکده منابع طبیعی، دانشگاه تربیت مدرس، تهران، ایران.
چکیده
چکیده

سابقه و هدف

فرونشست زمین (LS) در ایران به یکی از چالش‌های اصلی مدیریت منابع آب و خاک تبدیل شده است. در سال‌های اخیر، ترکیب داده‌های سنجش از دور با الگوریتم‌های یادگیری ماشین، رویکردی امیدبخش برای برآورد مکانی فرونشست ارائه داده است. لذا در این پژوهش سعی شده است کارایی همزمان ۹ الگوریتم پرکاربرد یادگیری ماشین در تعیین فرونشست در دشت‌های درگیر ایران تعیین شود. همچنین استفاده از روش تفسیرپذیری SHAP برای شناسایی سهم واقعی هر متغیر محیطی در خروجی مدل و تعیین مدل بهینه برای برآورد فرونشست با تأکید بر متغیر افت تراز آب زیرزمینی به عنوان مهم‌ترین عامل کنترل‌کنند از دیگر اهداف این پژوهش است.

مواد و روش‌ها:

در این پژوهش، از نقشه فرونشست سال ۲۰۲۰ ایران به عنوان داده مرجع استفاده گردید. تعداد تقریبا ۱7۰۰۰ نقطه نمونه به روش تصادفی طبقه‌بندی‌شده از این نقشه استخراج شد. برای هر نقطه، ۲۰ لایه اطلاعاتی شامل متغیرهای اقلیمی (بارش و دما از ۹۰ ایستگاه سینوپتیک)، توپوگرافیک (شاخص رطوبت توپوگرافی (TWI)، شیب، جهت شیب، عامل طول شیب، عمق دره، شاخص همواری دره(MrVBF)، هیدرولوژیک (افت تراز آب زیرزمینی)، انسانی (فاصله از جاده، گسل، رودخانه، مناطق مسکونی، اراضی کشاورزی، باغات)، و شاخص‌های طیفی مستخرج از تصاویر سنتینل-۲ (شامل NDVI، SAVI، EVI، BSI) تهیه گردید. همه لایه‌ها به قدرت تفکیک مکانی 150 متر و سیستم تصویر WGS84 هم‌مرجع شدند. سپس ۹ مدل یادگیری ماشین شامل جنگل تصادفی(RF)، ماشین بردار پشتیبان(SVM)، تقویت گرادیان شدید (XGBoost)، تقویت دسته‌ای (CatBoost)، درخت تصمیم(DT)، ماشین تقویت گرادیان سبک (LightGBM)، نزدیک‌ترین همسایه (kNN)، شبکه عصبی مصنوعی(ANN)، و درختان فوق تصادفی (ET) پیاده‌سازی شدند. داده‌ها به نسبت ۷۰ درصد آموزش، ۱۵ درصد اعتبارسنجی و ۱۵ درصد تست تقسیم شدند. برای ارزیابی عملکرد مدل‌ها از شاخص‌های ضریب تبیین (R²)، ریشه میانگین مربعات خطا (RMSE) و میانگین قدر مطلق خطا (MAE) استفاده گردید. همچنین، به دلیل تبدیل مسئله رگرسیون به دودویی (با آستانه فرونشست بحرانی ۵ سانتی‌متر در سال)، منحنی ROC و سطح زیرمنحنی AUC نیز محاسبه شدند. برای خارج کردن مدل از حالت جعبه سیاه و تعیین مشارکت نسبی هر متغیر، از روش SHAP مبتنی بر نظریه بازی‌های شپلی استفاده گردید.

نتایج و بحث:

نتایج نشان داد که مدل ET با ضریب تبیین ۰٫۹۹۲، کمترین میزان خطای RMSE برابر ۱٫۰۰۱ میلی‌متر و MAE برابر ۰٫۳۹۷، بالاترین دقت را در میان ۹ مدل دارد. این مدل با تمرکز بیشینه نقاط در امتداد خط ۱:۱ در نمودار پراکندگی و قرارگیری در باند خطای ۱۰ درصد، بیشترین انطباق را با مقادیر InSAR نشان داد. در مقابل، مدل ماشین بردار پشتیبان با ضریب تبیین منفی (۳۸۶/۰)، عملکرد بسیار ضعیفی داشت که به حساسیت بالای این مدل به مقیاس داده‌ها و نیاز به تنظیم دقیق فراپارامترها نسبت داده می‌شود. مدل‌های مبتنی بر جنگل تصادفی و تقویت گرادیان (CatBoost و XGBoost) با R² بالای ۰٫۹۸۷ در رتبه‌های بعدی قرار گرفتند. تحلیل منحنی ROC نشان داد مدل درختان فوق تصادفی با AUC بیش از ۰٫۹۰، بهترین توانایی را در تفکیک مناطق بحرانی از پایدار دارد. از نظر تفسیرپذیری، نمودارهای SHAP آشکار ساختند که در همه مدل‌ها و به ویژه در ET و CatBoost، متغیرهای «افت تراز آب زیرزمینی»، «تراکم جمعیت » و «اقلیم» بیش‌ترین سهم را در برآورد فرونشست داشته‌اند.

نتیجه‌گیری:

مدل ET به همراه تفسیرگر SHAP، به عنوان دقیق‌ترین و قابل‌اطمینان‌ترین رویکرد برای برآورد فرونشست در دشت‌های ایران معرفی می‌شود. استفاده از این مدل می‌تواند هزینه‌های پایش میدانی را کاهش داده و اولویت‌بندی مناطق بحرانی را برای مدیران منابع آب تسهیل کند. محدودیت اصلی پژوهش، عدم دسترسی به سری زمانی بلندمدت یکسان برای افت آب زیرزمینی در همه دشت‌ها است که پیشنهاد می‌شود در پژوهش‌های آتی با استفاده از مدل‌های هیدرولوژیک پویا جبران شود.

واژه‌های کلیدی: ایران، پایش فرونشست، تخریب منابع خاک و آب، هوش مصنوعی، مدیریت بهینه
کلیدواژه‌ها

عنوان مقاله English

Comparative evaluation of machine learning algorithms with an emphasis on Extra Trees and SHAP interpretability in estimating land subsidence in the plains of Iran

نویسندگان English

Karim Solaimani 1
fatemeh abedi 2
Reza Chamani 3
Fatemeh Akbari Emamzadeh 4
1 Director of Remote Sensing Center, Sari University
2 - PhD student, Watershed Science and Engineering, Faculty of Natural Resources, Sari University of Natural Resources, Sari, Iran.
3 - Assistant Professor, Rangeland and Watershed Management Department, Faculty of Natural Resources and Environment, University of Birjand, Birjand, Iran
4 PhD, Watershed Management Science and Engineering, Faculty of Natural Resources, Tarbiat Modares University, Tehran, Iran.
چکیده English

Abstract

Background and Objective

Land subsidence is one of the major challenges in water and soil resource management in Iran. In recent years, integrating remote sensing data with machine learning algorithms has provided a promising approach for the spatial estimation of subsidence. Therefore, this study aimed to evaluate the simultaneous performance of nine commonly used machine learning algorithms in identifying subsidence across the Iranian affected plains. Another objective was to employ the SHAP interpretability method to determine the actual contribution of each environmental variable to the model output and to identify the optimal model for subsidence estimation, with an emphasis on groundwater-level decline as the most controlling factor.

Materials and Methods

In this research, the land subsidence map of Iran for the year 2020 was used as the reference dataset. A total of 17,000 stratified random sample points were extracted from this map. For each point, 20 information layers were prepared, including climatic variables (precipitation and temperature from 90 synoptic stations), topographic indices (Topographic Wetness Index (TWI), slope, aspect, LS factor, valley depth, valley flatness index (MrVBF)), hydrological factors (groundwater-level decline), anthropogenic factors (distance from roads, faults, rivers, residential areas, agricultural lands, and orchards), and spectral indices derived from Sentinel-2 imagery (NDVI, SAVI, EVI, BSI). All layers were resampled to a spatial resolution of 150 meters and standardized to the WGS84 coordinate system. Subsequently, nine machine learning models were implemented, including Random Forest (RF), Support Vector Machine (SVM), Extreme Gradient Boosting (XGBoost), Categorical Boosting (CatBoost), Decision Tree (DT), Light Gradient Boosting Machine (LightGBM), k-Nearest Neighbors (kNN), Artificial Neural Network (ANN), and Extra Trees (ET). The data were split into training (70%), validation (15%), and testing (15%) subsets. Model performance was evaluated using the coefficient of determination (R²), root mean square error (RMSE), and mean absolute error (MAE). Furthermore, because the regression problem was converted into a binary classification using a critical subsidence threshold of 5 cm/year, ROC curves and AUC values were also calculated. To eliminate the "black box" nature of the models and determine the relative contribution of each variable, the SHAP method, based on Shapley game theory, was applied.

Results and Discussion

The results indicated that the ET model, with an R² of 0.992, the lowest RMSE of 1.001 mm, and an MAE of 0.397, demonstrated the highest accuracy among the nine models. This model showed the greatest alignment with InSAR-derived values, with the majority of data points concentrated along the 1:1 line in the scatter plot and falling within the 10% error band. In contrast, the SVM model exhibited very poor performance, with a negative R² (-0.386), which is attributed to its high sensitivity to data scaling and the need for precise hyperparameter tuning. Models based on Random Forest and gradient boosting (CatBoost and XGBoost) ranked next, with R² values above 0.987. ROC analysis confirmed that the Extra Trees model, with an AUC exceeding 0.90, had the best capability for distinguishing critical areas from stable ones. In terms of interpretability, SHAP plots revealed that across all models, and especially in ET and CatBoost, the variables "groundwater-level decline," "population density," and "climate" had the greatest contribution to subsidence estimation.

Conclusion

The Extra Trees model, combined with the SHAP explainer, is introduced as the most accurate and reliable approach for estimating subsidence across Iranian plains. The application of this framework can reduce field monitoring costs and facilitate the prioritization of critical zones for water resource managers. The main limitation of this research is the lack of consistent long-term temporal data on groundwater-level decline across all plains, which is recommended to be addressed in future studies using dynamic hydrological models.

Keywords: Iran, Subsidence monitoring, Soil and water resource degradation, Artificial Intelligence, Optimal management

کلیدواژه‌ها English

Keywords: Iran
Subsidence monitoring
Soil and water resource degradation
Artificial Intelligence
Optimal management

مقالات آماده انتشار، پذیرفته شده
انتشار آنلاین از 29 تیر 1405