A team is using Microsoft Foundry in a development environment to test and compare large language model (LLM) prompt variants.
The team needs consistent inputs to assess prompt variants without depending on live user traffic.
You need to create a controlled input-data evaluation. Which action should you take first?
Community Discussion