<p>Neuroimaging open-data initiatives have led to increased availability of large scientific datasets. While these datasets are shifting the processing bottleneck from compute-intensive to data-intensive, current standardized analysis tools have yet to adopt strategies that mitigate the costs associated with large data transfers. A major challenge in adapting neuroimaging applications for data-intensive processing is that they must be entirely rewritten. To facilitate data management for standardized neuroimaging tools, we developed Sea, a library that intercepts and redirects application read and write calls to minimize data transfer time. In this paper, we investigate the performance of Sea on three preprocessing pipelines applied to three different neuroimaging datasets on two high-performance computing clusters. Our results demonstrate that Sea provides large speedups (up to 32<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\times\)</EquationSource> </InlineEquation>) when the shared file system’s performance is deteriorated. When the shared file system is not overburdened by other users, performance is unaffected by Sea, suggesting that Sea’s overhead is minimal even in cases where its benefits are limited.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hierarchical Storage Management in User Space for Neuroimaging Applications

  • Valérie Hayot-Sasson,
  • Tristan Glatard

摘要

Neuroimaging open-data initiatives have led to increased availability of large scientific datasets. While these datasets are shifting the processing bottleneck from compute-intensive to data-intensive, current standardized analysis tools have yet to adopt strategies that mitigate the costs associated with large data transfers. A major challenge in adapting neuroimaging applications for data-intensive processing is that they must be entirely rewritten. To facilitate data management for standardized neuroimaging tools, we developed Sea, a library that intercepts and redirects application read and write calls to minimize data transfer time. In this paper, we investigate the performance of Sea on three preprocessing pipelines applied to three different neuroimaging datasets on two high-performance computing clusters. Our results demonstrate that Sea provides large speedups (up to 32 \(\times\) ) when the shared file system’s performance is deteriorated. When the shared file system is not overburdened by other users, performance is unaffected by Sea, suggesting that Sea’s overhead is minimal even in cases where its benefits are limited.