A Memristive Circuit Based on Q-Learning and Operant Conditioning
摘要
In most memristive neural network circuits based on operant conditioning, the agent’s tendency towards certain behaviors is simply reflected through changes in synaptic weight. No specific analysis has been conducted on the changes in the tendency of intelligent agent behavior. Therefore, an operant conditioning circuit based on memristors and Q-learning is proposed to analyze the behavior and decision-making of intelligent agents in complex environments. The designed network uses Q-learning to update the decision voltage in the circuit based on the optimal Bellman equation, allowing agents to make different strategies according to the constantly changing environment to achieve optimal results. In addition, the combination of Q-learning and memristors helps to alleviate the inherent overestimation bias in Q-learning, improve the stability and performance of the learning process. The proposed circuit can be applied in fields such as biomimetic robots, path planning and braking manufacturing, providing new ideas for the development of neuromorphic learning.