Autonomous agents: Augmenting visual information with raw audio data
Abstract
In the realm of game playing, deep reinforcement learning predominantly relies on visual input to map states to actions. The visual data extracted from the game environment serves as the primary foundation for state representation in reinforcement learning agents. However, humans leverage additional sensory inputs, such as audio cues, which play a pivotal role in perception and decision-making. Therefore, incorporating raw audio along with visual information shows potential for offering valuable insights to reinforcement learning agents. This study advocates for the integration of raw audio samples as complementary information to visual data in state representation. By using raw audio with visual cues, our objective is to enrich the decision-making process of the agent at each stage. Experimental evaluation were conducted employing Deep Q Networks (DQN) and Proximal Policy Optimization (PPO) algorithms within ViZDoom and Unity reinforcement learning environments. The results of our experiments reveal that augmenting visual information with raw audio samples yields superior rewards and expedites the learning rate compared to relying solely on visual data. Additionally, the findings suggest that considering both visual and audio features enhances the agent’s behavior, a trend observed across Unity and ViZDoom environments. This study underscores the potential advantages of incorporating multisensory information, particularly raw audio, into the state representation of reinforcement learning agents. Such insights contribute to advancing our understanding of how agents perceive and engage with their environments, ultimately enhancing performance in complex gaming scenarios.
Article Details
Authors (1)
Enoch Solomon