QuestionQ56
Operational Efficiency and Optimization for GenAI ApplicationsA financial services company uses an AI application to process financial documents with Amazon Bedrock. During business hours, the application processes approximately 10,000 requests per hour and therefore requires consistent throughput.
The company uses the CreateProvisionedModelThroughput API to buy provisioned throughput. Amazon CloudWatch metrics show that the provisioned capacity is unused while on-demand requests are being throttled. The company discovers the following code in the application:
response = bedrock_runtime.invoke_model(modelId="anthropic.claude-v2", body=json.dumps(payload))
The company needs the application to use the provisioned throughput and resolve the throttling issues.
Which solution will meet these requirements?
Community Discussion