QuestionQ7

AI Technologies and Controls

Which of the following would BEST help mitigate vulnerabilities related to hidden triggers in generative AI models?

Explanation

Adversarial training deliberately exposes a model to malicious or trigger-like inputs during training, allowing the model and its defenses to be evaluated and hardened against backdoor behavior. Hidden triggers are a form of backdoor/Trojan behavior in which a specific input pattern causes an attacker-chosen response.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!