<p>Healthcare data sharing is no longer an occasional exception; it has become an operational norm, and privacy risk has risen with it. Electronic health records are valuable precisely because they are information-dense: they support clinical and administrative analyses that would be impossible with thinner data. The catch is that this same richness often leaves “de-identified” releases more linkable than practitioners would like to admit. GUARDIAN is intended as a pragmatic answer to that tension. Designed as an integrated, healthcare-oriented privacy-preserving pipeline rather than a formal optimization solver, it coordinates adaptive <i>k</i>-anonymity with active <i>l</i>-diversity enforcement, <i>t</i>-closeness enforcement via distributional merging, and calibrated geospatial perturbation all within a configurable multi-stage architecture. On a synthetic Indian healthcare dataset of 10,000 patient records, the framework attains correlation preservation above 99% for claim amounts (99.9%) and length of stay (99.6%) alongside <i>k</i>-anonymity (<InlineEquation ID="IEq1"><EquationSource Format="TEX">\(k \ge 5\)</EquationSource></InlineEquation>), <i>l</i>-diversity (<InlineEquation ID="IEq2"><EquationSource Format="TEX">\(l \ge 3\)</EquationSource></InlineEquation>), <i>t</i>-closeness (mean EMD <InlineEquation ID="IEq3"><EquationSource Format="TEX">\(= 0.085\)</EquationSource></InlineEquation>, threshold 0.15), and geographic displacement (mean 4.79&#xa0;km). A matched-protection comparison where baselines are given the same proportional perturbation as GUARDIAN confirms that the utility advantage over conventional microaggregation baselines is attributable to the noise paradigm choice (Wilcoxon <InlineEquation ID="IEq4"><EquationSource Format="TEX">\(p = 0.002\)</EquationSource></InlineEquation>), while GUARDIAN’s genuine contribution lies in providing multi-layered privacy protections that single-mechanism baselines lack, at no additional utility cost (Wilcoxon <InlineEquation ID="IEq5"><EquationSource Format="TEX">\(p &gt; 0.14\)</EquationSource></InlineEquation>). Cross-dataset validation on an independently designed US healthcare dataset, multi-scale evaluation from 1000 to 100,000 records, and stratified subpopulation analysis across 16 demographic subgroups confirm robustness (bootstrap CV <InlineEquation ID="IEq6"><EquationSource Format="TEX">\(= 0.03\%\)</EquationSource></InlineEquation>), though all evaluations are conducted on synthetic data, and validation on real-world EHR datasets remains an important future direction. Under adversarial testing, GUARDIAN shows 82.2% resistance to record linkage attacks under full attacker knowledge using probabilistic Fellegi–Sunter scoring, suggesting that the protection is not fragile when assumptions are strengthened. A utility-aware feedback mechanism demonstrates adaptive noise recovery from 57.2 to 99.8% utility across 6 iterations, confirming that high utility is achievable through calibration rather than being an artifact of fixed parameters. Overall, the results indicate that under the reported synthetic evaluation setup, multi-layered privacy protections can be provided at no apparent additional utility cost compared to matched baselines, provided the mechanisms are coordinated and tuned with healthcare semantics in mind.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A multi-layered privacy-preserving framework for secure healthcare data sharing with high utility retention

  • Nagaraj S.,
  • Vijayarajan V.

摘要

Healthcare data sharing is no longer an occasional exception; it has become an operational norm, and privacy risk has risen with it. Electronic health records are valuable precisely because they are information-dense: they support clinical and administrative analyses that would be impossible with thinner data. The catch is that this same richness often leaves “de-identified” releases more linkable than practitioners would like to admit. GUARDIAN is intended as a pragmatic answer to that tension. Designed as an integrated, healthcare-oriented privacy-preserving pipeline rather than a formal optimization solver, it coordinates adaptive k-anonymity with active l-diversity enforcement, t-closeness enforcement via distributional merging, and calibrated geospatial perturbation all within a configurable multi-stage architecture. On a synthetic Indian healthcare dataset of 10,000 patient records, the framework attains correlation preservation above 99% for claim amounts (99.9%) and length of stay (99.6%) alongside k-anonymity (\(k \ge 5\)), l-diversity (\(l \ge 3\)), t-closeness (mean EMD \(= 0.085\), threshold 0.15), and geographic displacement (mean 4.79 km). A matched-protection comparison where baselines are given the same proportional perturbation as GUARDIAN confirms that the utility advantage over conventional microaggregation baselines is attributable to the noise paradigm choice (Wilcoxon \(p = 0.002\)), while GUARDIAN’s genuine contribution lies in providing multi-layered privacy protections that single-mechanism baselines lack, at no additional utility cost (Wilcoxon \(p > 0.14\)). Cross-dataset validation on an independently designed US healthcare dataset, multi-scale evaluation from 1000 to 100,000 records, and stratified subpopulation analysis across 16 demographic subgroups confirm robustness (bootstrap CV \(= 0.03\%\)), though all evaluations are conducted on synthetic data, and validation on real-world EHR datasets remains an important future direction. Under adversarial testing, GUARDIAN shows 82.2% resistance to record linkage attacks under full attacker knowledge using probabilistic Fellegi–Sunter scoring, suggesting that the protection is not fragile when assumptions are strengthened. A utility-aware feedback mechanism demonstrates adaptive noise recovery from 57.2 to 99.8% utility across 6 iterations, confirming that high utility is achievable through calibration rather than being an artifact of fixed parameters. Overall, the results indicate that under the reported synthetic evaluation setup, multi-layered privacy protections can be provided at no apparent additional utility cost compared to matched baselines, provided the mechanisms are coordinated and tuned with healthcare semantics in mind.