A Copilot Studio agent is assessed with a fixed test set and an automated evaluation method.
After the evaluation runs, the team finds that the same three interactions fail on every run, whereas all other interactions pass consistently.
You need to use the evaluation results to identify an accurate conclusion.
What should you conclude from the evaluation results?
Community Discussion