Integrating Habituation Effects with UCB and Softmax Multi-Armed Bandit Algorithms for Optimized Digital Content Delivery
摘要
In the realm of Multi-Armed Bandits (MAB), the integration of habituation with intermittent breaks emerges as a compelling paradigm to enhance exploration-exploitation trade-offs. This study thoroughly investigates the application of habituation with breaks in two prominentstrategies: Softmax and Upper Confidence Bound (UCB). Empirical findings indicate that habituation with breaks can enhance long-term performance, mitigate reward stagnation, and support continuous adaptation. By elucidating the nuanced interplay between habituation, breaks, and MAB strategies, this work aims to inform future developments in decision-making algorithms designed for dynamic, real-world applications.