<p>Efficient resource selection and scheduling in large-scale computing systems like the Cloud pose significant challenges, especially for workflows with unreliable or preemptible resources. These workflows depict a set of (sequential or parallel) processing activities that should be executed upon the available resources. In this paper, we focus on the efficient management of resources paying attention to unreliable executions and propose a novel checkpointing mechanism integrated into an uncertainty-aware Heterogeneous Earliest Finish Time (uHEFT) online algorithm. This algorithm determines the optimal time instances to store the observed progress of the execution by adopting the principles of the Optimal Stopping Theory and the continuous monitoring of the execution parameters. Hence, we are able to monitor the execution of tasks on unreliable resources and determine the optimal points to store progress, thereby mitigating the risk of unexpected revocations which frequently occur in preemptive strategies. The proposed model leverages the heterogeneity of available Cloud services, providing a proactive mitigation approach for revocation risks. The pros and cons of our approach are evaluated through a set of extensive simulations and performance metrics against an optimal strategy in which every checkpoint is applied just before a revocation event is observed. Notably, uHEFT minimizes the deviation from the theoretical optimal strategy (regret) without increasing the overall number of checkpoints. This cost-efficient solution enhances the execution of workflows in environments with volatile resources.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A proactive and uncertainty driven management mechanism for workflows of processing tasks

  • Panagiotis Oikonomou,
  • Kostas Kolomvatsos,
  • Christos Anagnostopoulos

摘要

Efficient resource selection and scheduling in large-scale computing systems like the Cloud pose significant challenges, especially for workflows with unreliable or preemptible resources. These workflows depict a set of (sequential or parallel) processing activities that should be executed upon the available resources. In this paper, we focus on the efficient management of resources paying attention to unreliable executions and propose a novel checkpointing mechanism integrated into an uncertainty-aware Heterogeneous Earliest Finish Time (uHEFT) online algorithm. This algorithm determines the optimal time instances to store the observed progress of the execution by adopting the principles of the Optimal Stopping Theory and the continuous monitoring of the execution parameters. Hence, we are able to monitor the execution of tasks on unreliable resources and determine the optimal points to store progress, thereby mitigating the risk of unexpected revocations which frequently occur in preemptive strategies. The proposed model leverages the heterogeneity of available Cloud services, providing a proactive mitigation approach for revocation risks. The pros and cons of our approach are evaluated through a set of extensive simulations and performance metrics against an optimal strategy in which every checkpoint is applied just before a revocation event is observed. Notably, uHEFT minimizes the deviation from the theoretical optimal strategy (regret) without increasing the overall number of checkpoints. This cost-efficient solution enhances the execution of workflows in environments with volatile resources.