In this work, we explore the Transcriptomics Atlas pipeline adapted for cost-efficient and high-throughput computing in the cloud. We propose a scalable, cloud-native architecture designed for running a resource-intensive aligner – STAR – and processing hundreds of terabytes of RNA-sequencing data. We implement optimization techniques that significantly reduce cost and execution time. The impact of particular optimizations is measured in medium-scale experiments followed by a large-scale experiment that leverages all of them and validates the design. Early stopping optimization allows us to reduce the total alignment time by 23%. For the cloud environment, we identify suitable EC2 instance types and verify the applicability of spot instances usage.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Accelerating Cloud-Based Transcriptomics: Performance Analysis and Optimization of the STAR Aligner Workflow

  • Piotr Kica,
  • Sabina Lichołai,
  • Michał Orzechowski,
  • Maciej Malawski

摘要

In this work, we explore the Transcriptomics Atlas pipeline adapted for cost-efficient and high-throughput computing in the cloud. We propose a scalable, cloud-native architecture designed for running a resource-intensive aligner – STAR – and processing hundreds of terabytes of RNA-sequencing data. We implement optimization techniques that significantly reduce cost and execution time. The impact of particular optimizations is measured in medium-scale experiments followed by a large-scale experiment that leverages all of them and validates the design. Early stopping optimization allows us to reduce the total alignment time by 23%. For the cloud environment, we identify suitable EC2 instance types and verify the applicability of spot instances usage.