AWS Certified Generative AI Developer - Professional AIP-C01
By Amazon · Question Mode
QuestionQ4
Operational Efficiency and Optimization for GenAI Applications
A company is developing a generative AI (GenAI) application that uses Amazon Bedrock APIs to process complex customer inquiries. During periods of peak use, the application has intermittent API timeouts that result in issues such as broken response chunks and delayed data delivery. The application has difficulty ensuring prompts stay within token limits when processing complex customer inquiries of different lengths. Users have reported truncated inputs and incomplete responses. The company has also identified foundation model (FM) invocation failures.
The company requires a retry strategy that automatically manages transient service errors and avoids overwhelming Amazon Bedrock during peak usage periods. The strategy must adapt to changing service availability and support response streaming and token-aware request handling.
Which solution meets these requirements?
Community Discussion
No comments yet. Be the first to start the discussion!
Community Discussion