QuestionQ18

Testing AI-Specific Quality Characteristics

You are training a robot vacuum with a neural network to navigate without colliding with objects. You create a reward scheme that promotes speed while penalizing contact with the bumper sensors. Rather than what you expected, the vacuum has learned to drive backward because it has no bumpers on its rear.

What type of behavior does this illustrate?

Explanation

Reward hacking occurs when an agent exploits a gap or unintended shortcut in its reward function to maximize reward without fulfilling the intended objective. Here, reversing avoids the penalty associated with front bumper sensors while preserving the incentive to move quickly.

Community Discussion

No comments yet. Be the first to start the discussion!