水下物联网中基于深度强化学习算法的中继选择方案

RELAY SELECTION SCHEME BASED ON DEEP REINFORCEMENT LEARNING IN INTERNET OF UNDERWATER THINGS

  • 摘要: 针对水下物联网(Internet of Underwater Things,IoUT)通信覆盖和能量资源的受限问题,提出一种结合了最佳发射功率与近端策略优化(Proximal Policy Optimization,PPO)算法的水下中继选择策略。策略建立水下协作中继过程的马尔可夫模型,通过凸优化的方法确定源节点和所选水下中继节点的最佳发射功率,最大化信噪比。仿真结果表明,相同条件下,与随机中继选择和Q学习(Q-learning,QL)算法进行对比该方案具有更高的累积奖励、更高的信噪比、更低的中断概率。

     

    Abstract: To address the issues of limited communication coverage and energy resources in the Underwater Internet of Things (IoUT), a relay selection strategy is proposed that combines optimal transmission power with proximal policy optimization (PPO). The strategy established a Markov model for the underwater cooperative relay process, and employed convex optimization to determine the optimal transmission power for the source node and the selected underwater relay node, maximizing the signal- to- noise ratio. Simulation results show that under the same conditions, compared with random relay selection and Q- learning (QL) algorithm, this approach has higher cumulative rewards, a higher signal- to- noise ratio, and lower interruption probability.

     

/

返回文章
返回