<p>Many organizations are producing or collecting private data in fields such as medical research and government regulation. Due to current privacy protection, laws and regulations, commercial competition, and other issues, these institutions cannot directly share their data. Collaborative analysis of private data from multiple institutions will benefit each institution and create profits together. Therefore, we propose a Yannakakis-based multiparty outsourcing collaboration analysis scheme. It enables organizations to collaboratively analyze private data from multiple organizations according to their needs while ensuring that private data are not leaked to each other. Our scheme is based on the improved Yannakakis algorithm to build a series of query components, such as Semi-join, Join, Order-by, etc. We also optimized the join operation. By confusing the input tuples and protecting their authenticity through annotations, the join operation can be directly joined through the hash value without disclosing the join results. Through this series of configurations, we can execute a query with <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_6994_Article_IEq1.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="102" /> </InlineMediaObject> <EquationSource Format="TEX">\( O(\textrm{IN}+\textrm{OUT}) \)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>O</mi> <mo stretchy="false">(</mo> <mtext>IN</mtext> <mo>+</mo> <mtext>OUT</mtext> <mo stretchy="false">)</mo> </mrow> </math></EquationSource> </InlineEquation> runtime and communication, where <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_6994_Article_IEq2.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="20" /> </InlineMediaObject> <EquationSource Format="TEX">\( \textrm{IN} \)</EquationSource> <EquationSource Format="MATHML"><math> <mtext>IN</mtext> </math></EquationSource> </InlineEquation> is the total number of tuples in the input relationship, and <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_6994_Article_IEq3.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="38" /> </InlineMediaObject> <EquationSource Format="TEX">\( \textrm{OUT} \)</EquationSource> <EquationSource Format="MATHML"><math> <mtext>OUT</mtext> </math></EquationSource> </InlineEquation> is the output size. We have carried out a series of comparative experiments, and the results show that our system is 1.3<i>X</i>–7.4<i>X</i> faster than the baseline.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Outsourcing collaboration analysis of multiparty privacy data using the improved Yannakakis

  • Zigang Chen,
  • Zhenjiang Zhang,
  • Tao Leng,
  • Haihua Zhu,
  • Yuhong Liu

摘要

Many organizations are producing or collecting private data in fields such as medical research and government regulation. Due to current privacy protection, laws and regulations, commercial competition, and other issues, these institutions cannot directly share their data. Collaborative analysis of private data from multiple institutions will benefit each institution and create profits together. Therefore, we propose a Yannakakis-based multiparty outsourcing collaboration analysis scheme. It enables organizations to collaboratively analyze private data from multiple organizations according to their needs while ensuring that private data are not leaked to each other. Our scheme is based on the improved Yannakakis algorithm to build a series of query components, such as Semi-join, Join, Order-by, etc. We also optimized the join operation. By confusing the input tuples and protecting their authenticity through annotations, the join operation can be directly joined through the hash value without disclosing the join results. Through this series of configurations, we can execute a query with \( O(\textrm{IN}+\textrm{OUT}) \) O ( IN + OUT ) runtime and communication, where \( \textrm{IN} \) IN is the total number of tuples in the input relationship, and \( \textrm{OUT} \) OUT is the output size. We have carried out a series of comparative experiments, and the results show that our system is 1.3X–7.4X faster than the baseline.