<p>The early detection of periodontitis poses an ongoing challenge in clinical practice due to the limitations of traditional diagnostic methods. While new AI-powered technologies offer promising solutions, the lack of high-quality, publicly available biomarker datasets continues to hinder progress. This study explores the use of synthetic data generation, an under-utilized approach in periodontal research but gaining traction in other medical fields, to support the development of machine learning models for distinguishing health states related to periodontitis. A synthetic dataset of 4000 samples was generated using a Tabular Variational Autoencoder, incorporating inflammatory, metabolic, microbiome, and demographic features. Three health states were modeled: Healthy, Periodontitis, and an intermediate Elevated class designed to reflect early-stage disease. Several ML models, including XGBoost, Random Forest, and Logistic Regression, were evaluated, with XGBoost achieving the highest performance (accuracy: 84%; F1-score: 81%). Feature attribution analyses identified IL-1β, IL-10, and urea as influential predictors, illustrating the potential of synthetic data to reveal model–biomarker interactions. This work demonstrates the potential of synthetic data as a complementary tool for AI-driven periodontal diagnostics, particularly for exploring early-stage detection strategies. By capturing biomarker patterns across health states, synthetic datasets provide a foundation for model development and hypothesis generation when real-world data are limited. The present findings should be viewed as proof of concept, since both training and evaluation relied solely on synthetic data without clinical validation. Empirical validation remains essential for translating these findings into clinical applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Synthetic data as a tool for prototyping early-stage periodontitis detection models

  • Erdal Akin,
  • Filip Kroon,
  • Cassandra Windahl,
  • Yutaka Sugihara,
  • Aleksandar Milosavljevic,
  • Magnus Falk

摘要

The early detection of periodontitis poses an ongoing challenge in clinical practice due to the limitations of traditional diagnostic methods. While new AI-powered technologies offer promising solutions, the lack of high-quality, publicly available biomarker datasets continues to hinder progress. This study explores the use of synthetic data generation, an under-utilized approach in periodontal research but gaining traction in other medical fields, to support the development of machine learning models for distinguishing health states related to periodontitis. A synthetic dataset of 4000 samples was generated using a Tabular Variational Autoencoder, incorporating inflammatory, metabolic, microbiome, and demographic features. Three health states were modeled: Healthy, Periodontitis, and an intermediate Elevated class designed to reflect early-stage disease. Several ML models, including XGBoost, Random Forest, and Logistic Regression, were evaluated, with XGBoost achieving the highest performance (accuracy: 84%; F1-score: 81%). Feature attribution analyses identified IL-1β, IL-10, and urea as influential predictors, illustrating the potential of synthetic data to reveal model–biomarker interactions. This work demonstrates the potential of synthetic data as a complementary tool for AI-driven periodontal diagnostics, particularly for exploring early-stage detection strategies. By capturing biomarker patterns across health states, synthetic datasets provide a foundation for model development and hypothesis generation when real-world data are limited. The present findings should be viewed as proof of concept, since both training and evaluation relied solely on synthetic data without clinical validation. Empirical validation remains essential for translating these findings into clinical applications.