QuestionQ56

Operational Efficiency and Optimization for GenAI Applications

A financial services company uses an AI application to process financial documents with Amazon Bedrock. During business hours, the application processes approximately 10,000 requests per hour and therefore requires consistent throughput.

The company uses the CreateProvisionedModelThroughput API to buy provisioned throughput. Amazon CloudWatch metrics show that the provisioned capacity is unused while on-demand requests are being throttled. The company discovers the following code in the application:

response = bedrock_runtime.invoke_model(modelId="anthropic.claude-v2", body=json.dumps(payload))  

The company needs the application to use the provisioned throughput and resolve the throttling issues.

Which solution will meet these requirements?

Explanation

Amazon Bedrock routes inference to Provisioned Throughput only when the runtime request supplies that provisioned model's ARN as the modelId. Supplying the foundation-model ID invokes on-demand capacity instead, leaving purchased provisioned capacity unused and subjecting the workload to on-demand throttling. Use a Provisioned Throughput with an Amazon Bedrock resource

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!