Advanced analysis of algorithms in multi-armed slot machine problems

Authors

  • Run Bai

DOI:

https://doi.org/10.54097/ht9yp521

Keywords:

Advanced analysis; Multi arm slot machine problems.

Abstract

The primary goal of the Multi-Armed Bandit (MAB) framework is to strike an optimal balance between exploring new strategies and exploiting existing ones. This framework, implemented through sophisticated algorithms, is designed to maximize benefits while minimizing resource expenditure. It finds widespread application across various fields, notably by data analysts, medical researchers, and marketing specialists. In the realm of data analysis, MAB algorithms are pivotal in optimizing choices based on evolving data trends. Medical researchers leverage these algorithms to make informed decisions during clinical trials, ensuring effective resource allocation and patient treatment strategies. In the marketing sector, MAB is instrumental in tailoring strategies to consumer behavior, enhancing customer engagement and improving marketing campaign efficiency. Overall, the MAB framework has established itself as an indispensable tool in decision-making processes, particularly in environments fraught with uncertainty. Its ability to adaptively allocate resources based on continuous learning and feedback makes it a cornerstone in numerous industries seeking to navigate complex and dynamic challenges.

Downloads

Download data is not yet available.

References

Aziz, M., Kaufmann, E., & Riviere, M. K. (2021). On multi-armed bandit designs for dose-finding clinical trials. The Journal of Machine Learning Research, 22(1), 686-723.

Yang, Z., Liu, X., & Ying, L. (2022, September). Exploration. Exploitation, and Engagement in Multi-Armed Bandits with Abandonment. In 2022 58th Annual Allerton Conference on Communication, Control, and Computing (Allerton) (pp. 1-2). IEEE.

Yang, Z., Liu, X., & Ying, L. (2022, September). Exploration. Exploitation, and Engagement in Multi-Armed Bandits with Abandonment. In 2022 58th Annual Allerton Conference on Communication, Control, and Computing (Allerton) (pp. 1-2). IEEE.

Zhu, X., Huang, Y., Wang, X., & Wang, R. (2023). Emotion recognition based on brain-like multimodal hierarchical perception. Multimedia Tools and Applications, 1-19.

Lai, T. L., & Robbins, H. (1985). Asymptotically efficient adaptive allocation rules. Advances in applied mathematics, 6(1), 4-22.

Vermorel, J., & Mohri, M. (2005, October). Multi-armed bandit algorithms and empirical evaluation. In European conference on machine learning (pp. 437-448). Berlin, Heidelberg: Springer Berlin Heidelberg.

Watkins, C. J. C. H. (1989). Learning from delayed rewards.

Lattimore, T., & Szepesvari, C. (2020). Bandit Algorithms. Cambridge University Press.

Wang, P. A., Proutiere, A., Ariu, K., Jedra, Y., & Russo, A. (2020, June). Optimal algorithms for multiplayer multi-armed bandits. In International Conference on Artificial Intelligence and Statistics (pp. 4120-4129). PMLR.

Liu, C. Y., & Li, L. (2016, September). On the prior sensitivity of thompson sampling. In International Conference on Algorithmic Learning Theory (pp. 321-336). Cham: Springer International Publishing.

Downloads

Published

26-04-2024

How to Cite

Bai, R. (2024). Advanced analysis of algorithms in multi-armed slot machine problems . Highlights in Science, Engineering and Technology, 94, 285-288. https://doi.org/10.54097/ht9yp521