Research on Sepsis using big data technologies often requires complex queries to explore the data beforehand. However, when dealing with massive datasets, research efficiency can be hindered by the frequent and time-consuming complex queries. The efficiency of these complex queries can be improved through index techniques. Compared to traditional indexes, learned indexes that incorporate deep learning have been proven to perform better. The two key aspects of learned indexes are the approximation of the cumulative distribution function (CDF) and the design of the index structure. Directly approximating the CDF on the original distribution is relatively inefficient, but converting a complex distribution into a simpler one beforehand can make the approximation of CDF more effective. Therefore, we propose a lightweight index structure based on data distribution transformation (DTLI). Firstly, we transform the original complex distribution into a smoother one through normalization flow. Then, we approximate the CDF on the transformed distribution using mutation points. Finally, we propose a lightweight learned index structure to use. Experimental results demonstrate that DTLI achieves higher throughput and lower latency in complex queries compared to traditional indexes and other learned indexes on the Sepsis queries.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Complex Query Optimization in Sepsis Data Using Learned Index with Distribution Transformation

  • Chao Xu,
  • Ye Liang,
  • Kaijun Wen,
  • Zihang Wang,
  • Yining Zhou,
  • Yong Zhang

摘要

Research on Sepsis using big data technologies often requires complex queries to explore the data beforehand. However, when dealing with massive datasets, research efficiency can be hindered by the frequent and time-consuming complex queries. The efficiency of these complex queries can be improved through index techniques. Compared to traditional indexes, learned indexes that incorporate deep learning have been proven to perform better. The two key aspects of learned indexes are the approximation of the cumulative distribution function (CDF) and the design of the index structure. Directly approximating the CDF on the original distribution is relatively inefficient, but converting a complex distribution into a simpler one beforehand can make the approximation of CDF more effective. Therefore, we propose a lightweight index structure based on data distribution transformation (DTLI). Firstly, we transform the original complex distribution into a smoother one through normalization flow. Then, we approximate the CDF on the transformed distribution using mutation points. Finally, we propose a lightweight learned index structure to use. Experimental results demonstrate that DTLI achieves higher throughput and lower latency in complex queries compared to traditional indexes and other learned indexes on the Sepsis queries.