QuestionQ19

Optimize generative AI systems and model performance

An organization is deploying generative AI solutions by using Microsoft Foundry to support multiple production workloads.

The organization has these workload requirements:

  • One workload must be real-time, latency-sensitive, and have predictable global usage patterns that require consistent performance.
  • One workload must have variable performance and be optimized for cost-efficient operation.

You need to choose a global deployment type for each workload.

Each deployment type may be used once, more than once, or not at all.

Drag & Drop
Global Provisioned
Global Standard
Real-time, predictable
Flexible, cost efficient
Explanation

Global Provisioned uses reserved capacity to deliver predictable throughput and lower latency variance, making it appropriate for real-time, latency-sensitive workloads with consistent demand. Global Standard uses pay-per-token billing and is suited to variable or bursty traffic, which makes it the more cost-efficient option when demand varies.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!