<p>Berlin has historically been impacted by heavy metal (HM) emissions, raising concerns about soil pollution. In this study, machine learning (ML) techniques were applied to predict HM concentrations across the Berlin metropolitan area. A dataset of 667 soil samples was used, containing spectrometry data and concentrations of nine HMs: arsenic&#xa0;(As), cadmium (Cd), cobalt (Co), chromium (Cr), copper (Cu), manganese (Mn), nickel (Ni), lead (Pb), and zinc (Zn). Four ML algorithms, including partial least square regression (PLSR), support vector machine regression (SVMR), random forest (RF), and Gaussian process regression (GPR), were employed to model and predict HMs. To address the often-ignored spatial dimension, samples were also stratified into five land use/land cover (LULC) classes: park, forest, farmland, traffic area, and constructed land. Among the full dataset, SVMR yielded the best performance in predicting Zn (<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10661_2025_14378_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="24" /> </InlineMediaObject> <EquationSource Format="TEX">\(\varvec{R^2}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msup> <mi mathvariant="bold-italic">R</mi> <mn mathvariant="bold">2</mn> </msup> </mrow> </math></EquationSource> </InlineEquation> = 0.65, RMSE = 34.61 mg/kg, ME = 0.40). For stratified modeling, PLSR performed best for Ni in farmland (<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10661_2025_14378_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="24" /> </InlineMediaObject> <EquationSource Format="TEX">\(\varvec{R^2}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msup> <mi mathvariant="bold-italic">R</mi> <mn mathvariant="bold">2</mn> </msup> </mrow> </math></EquationSource> </InlineEquation> = 0.77), while RF was most accurate for Ni in traffic areas (<InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="10661_2025_14378_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="24" /> </InlineMediaObject> <EquationSource Format="TEX">\(\varvec{R^2}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msup> <mi mathvariant="bold-italic">R</mi> <mn mathvariant="bold">2</mn> </msup> </mrow> </math></EquationSource> </InlineEquation> = 0.82). The results highlighted improved model performance within farmland and traffic areas compared to the unstratified dataset, demonstrating the value of incorporating spatial context in soil HM prediction.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unlocking urban soil secrets: machine learning and spectrometry in Berlin’s heavy metal pollution study considering spatial data

  • Mahsa Nakhostin Rouhi,
  • Mohammadmehdi Saberioon,
  • Mohsen Makki,
  • Saeid Homayouni,
  • Majid Kiavarz,
  • Seyed Kazem Alavipanah

摘要

Berlin has historically been impacted by heavy metal (HM) emissions, raising concerns about soil pollution. In this study, machine learning (ML) techniques were applied to predict HM concentrations across the Berlin metropolitan area. A dataset of 667 soil samples was used, containing spectrometry data and concentrations of nine HMs: arsenic (As), cadmium (Cd), cobalt (Co), chromium (Cr), copper (Cu), manganese (Mn), nickel (Ni), lead (Pb), and zinc (Zn). Four ML algorithms, including partial least square regression (PLSR), support vector machine regression (SVMR), random forest (RF), and Gaussian process regression (GPR), were employed to model and predict HMs. To address the often-ignored spatial dimension, samples were also stratified into five land use/land cover (LULC) classes: park, forest, farmland, traffic area, and constructed land. Among the full dataset, SVMR yielded the best performance in predicting Zn ( \(\varvec{R^2}\) R 2 = 0.65, RMSE = 34.61 mg/kg, ME = 0.40). For stratified modeling, PLSR performed best for Ni in farmland ( \(\varvec{R^2}\) R 2 = 0.77), while RF was most accurate for Ni in traffic areas ( \(\varvec{R^2}\) R 2 = 0.82). The results highlighted improved model performance within farmland and traffic areas compared to the unstratified dataset, demonstrating the value of incorporating spatial context in soil HM prediction.