Background <p>Programmed death-ligand 1 (PD-L1) expression and tumor mutational burden (TMB) are widely used immunotherapy biomarkers, yet their variability and determinants in lung squamous cell carcinoma (LUSC) remain incompletely characterized in real-world testing settings.</p> Methods <p>We analyzed a retrospective real-world LUSC cohort with clinical PD-L1 immunohistochemistry and next-generation sequencing. PD-L1 tumor proportion score (TPS) was modeled as an ordered endpoint using three clinically used categories (&lt; 1%, 1–49%, ≥ 50%) via proportional-odds ordinal regression; TMB was modeled as a continuous outcome after log transformation using multivariable linear regression. To address the practical question of whether imperfect real-world biomarker data can still support hypothesis generation, we additionally applied random forest and XGBoost as exploratory tools for prediction-oriented feature prioritization and evaluated uncertainty using repeated resampling and permutation-based perturbation analyses.</p> Results <p>Regression models provided an interpretable adjusted effect landscape but yielded limited statistically significant associations after adjustment. In contrast, tree-based models prioritized a small subset of clinicogenomic features with predictive information for PD-L1 TPS and TMB variability. Resampling and permutation analyses provided an internal assessment of uncertainty, rank stability, and sensitivity of these feature-prioritization signals.</p> Conclusions <p>In small, heterogeneous real-world LUSC cohorts where conventional regression-based inference may provide limited resolution, machine learning–guided feature prioritization combined with resampling- and permutation-based uncertainty assessment may help identify candidate clinicogenomic signals associated with PD-L1 TPS category and TMB variability. These findings should be interpreted as hypothesis-generating and require validation in larger independent cohorts.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine learning–guided identification of determinants of PD-L1 expression and tumor mutational burden in lung squamous cell carcinoma

  • Hongwei Shang,
  • Chao Li,
  • Jundong Wang,
  • Dongyun Xu,
  • Kailun Bai,
  • Shuiqing Zhou

摘要

Background

Programmed death-ligand 1 (PD-L1) expression and tumor mutational burden (TMB) are widely used immunotherapy biomarkers, yet their variability and determinants in lung squamous cell carcinoma (LUSC) remain incompletely characterized in real-world testing settings.

Methods

We analyzed a retrospective real-world LUSC cohort with clinical PD-L1 immunohistochemistry and next-generation sequencing. PD-L1 tumor proportion score (TPS) was modeled as an ordered endpoint using three clinically used categories (< 1%, 1–49%, ≥ 50%) via proportional-odds ordinal regression; TMB was modeled as a continuous outcome after log transformation using multivariable linear regression. To address the practical question of whether imperfect real-world biomarker data can still support hypothesis generation, we additionally applied random forest and XGBoost as exploratory tools for prediction-oriented feature prioritization and evaluated uncertainty using repeated resampling and permutation-based perturbation analyses.

Results

Regression models provided an interpretable adjusted effect landscape but yielded limited statistically significant associations after adjustment. In contrast, tree-based models prioritized a small subset of clinicogenomic features with predictive information for PD-L1 TPS and TMB variability. Resampling and permutation analyses provided an internal assessment of uncertainty, rank stability, and sensitivity of these feature-prioritization signals.

Conclusions

In small, heterogeneous real-world LUSC cohorts where conventional regression-based inference may provide limited resolution, machine learning–guided feature prioritization combined with resampling- and permutation-based uncertainty assessment may help identify candidate clinicogenomic signals associated with PD-L1 TPS category and TMB variability. These findings should be interpreted as hypothesis-generating and require validation in larger independent cohorts.