Purpose <p>Deep learning can classify histopathological subtypes of non-small-cell lung cancer (NSCLC) but often requires large datasets and heavy compute, limiting clinical use. We evaluate a lightweight alternative that pairs multi-domain hand-engineered features with conventional machine-learning classifiers.</p> Methods <p>We assembled a moderate dataset of 644 NSCLC histopathology images and extracted diverse features spanning original histopathological descriptors (e.g., keratinization field, glandular differentiation), physics-based measures (entropy, fractal dimension), mathematical descriptors (Hu moments, DCT/FFT energy), aesthetics (colorfulness, prevailing color), histologist-relevant metrics (nuclear perimeter/area), statistical summaries (percentiles, dispersion), and computer-vision features (HOG, keypoints). Random-forest–based feature selection identified informative predictors used to train support vector machines (SVM), gradient boosting, logistic regression, and decision trees. Performance was assessed via accuracy, recall, F1-score, and AUC with five-fold cross-validation. Generalizability was tested on an external set of 95 NSCLC images.</p> Results <p>Multi-domain features achieved strong performance with classical models. SVM obtained the highest cross-validated accuracy (0.85) and mean AUC (0.89). Gradient boosting was competitive (accuracy 0.84; AUC 0.87). Logistic regression outperformed decision trees overall. On external validation, the SVM trained with multi-domain features maintained robust performance (accuracy 0.84; AUC 0.85), indicating good generalizability despite the moderate training set.</p> Conclusion <p>SVM classifiers trained on carefully selected multi-domain features provide stable, high-quality NSCLC subtype classification without reliance on GPUs or large datasets. This approach offers a practical, resource-efficient pathway for clinical deployment, where computational capacity and data volume are often constrained.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-domain Histopathological Imaging Biomarkers for Machine-learning-based Classification of Lung Squamous Cell Carcinoma and Adenocarcinoma

  • Shilong Liu,
  • Jianli Ma,
  • Ke Jin,
  • Yixuan Wang,
  • Xueqiao Yu,
  • Kexin Chen,
  • Wencheng Shao

摘要

Purpose

Deep learning can classify histopathological subtypes of non-small-cell lung cancer (NSCLC) but often requires large datasets and heavy compute, limiting clinical use. We evaluate a lightweight alternative that pairs multi-domain hand-engineered features with conventional machine-learning classifiers.

Methods

We assembled a moderate dataset of 644 NSCLC histopathology images and extracted diverse features spanning original histopathological descriptors (e.g., keratinization field, glandular differentiation), physics-based measures (entropy, fractal dimension), mathematical descriptors (Hu moments, DCT/FFT energy), aesthetics (colorfulness, prevailing color), histologist-relevant metrics (nuclear perimeter/area), statistical summaries (percentiles, dispersion), and computer-vision features (HOG, keypoints). Random-forest–based feature selection identified informative predictors used to train support vector machines (SVM), gradient boosting, logistic regression, and decision trees. Performance was assessed via accuracy, recall, F1-score, and AUC with five-fold cross-validation. Generalizability was tested on an external set of 95 NSCLC images.

Results

Multi-domain features achieved strong performance with classical models. SVM obtained the highest cross-validated accuracy (0.85) and mean AUC (0.89). Gradient boosting was competitive (accuracy 0.84; AUC 0.87). Logistic regression outperformed decision trees overall. On external validation, the SVM trained with multi-domain features maintained robust performance (accuracy 0.84; AUC 0.85), indicating good generalizability despite the moderate training set.

Conclusion

SVM classifiers trained on carefully selected multi-domain features provide stable, high-quality NSCLC subtype classification without reliance on GPUs or large datasets. This approach offers a practical, resource-efficient pathway for clinical deployment, where computational capacity and data volume are often constrained.