Journal
IEEE CONFERENCE ON COMPUTER COMMUNICATIONS (IEEE INFOCOM 2019)
Volume -, Issue -, Pages 1693-1701Publisher
IEEE
DOI: 10.1109/infocom.2019.8737653
Keywords
Multi-player Multi-armed Bandits; Optimal regret; Pure exploration; Distributed learning
Categories
Funding
- Department of Science and Technology, India [IFA-14/ENG-73]
- Bharti Centre for Communication
- IIT Bombay [16IRCCSG010]
Ask authors/readers for more resources
We consider an ad hoc network where multiple users access the same set of channels. The channel characteristics are unknown and could be different tiff each user (heterogeneous). No controller is available to coordinate channel selections by the users, and if multiple users select the same channel, they collide and none of them receive any rate (or reward). For such a completely decentralized network we develop algorithms that aim to achieve optimal network throughput. Due to lack of any direct communication between the users, we allow each user to exchange information by transmitting in a specific pattern and sense such transmissions from others. However, such transmissions and sensing for information exchange do not add to network throughput. For the wideband sensing and narrowband sensing scenarios, we first develop explore-and-commit algorithms that converge to near-optimal allocation with high probability in a small number of rounds. Building on this, we develop an algorithm that gives logarithmic regret. We validate our claims through extensive experiments and show that our algorithms perform significantly better than the state-of-the-art CSM-MAB, dE(3) and dE(3)-TS algorithms.
Authors
I am an author on this paper
Click your name to claim this paper and add it to your profile.
Reviews
Recommended
No Data Available