QuestionQ4

Operational Efficiency and Optimization for GenAI Applications

A company is developing a generative AI (GenAI) application that uses Amazon Bedrock APIs to process complex customer inquiries. During periods of peak use, the application has intermittent API timeouts that result in issues such as broken response chunks and delayed data delivery. The application has difficulty ensuring prompts stay within token limits when processing complex customer inquiries of different lengths. Users have reported truncated inputs and incomplete responses. The company has also identified foundation model (FM) invocation failures.

The company requires a retry strategy that automatically manages transient service errors and avoids overwhelming Amazon Bedrock during peak usage periods. The strategy must adapt to changing service availability and support response streaming and token-aware request handling.

Which solution meets these requirements?

Explanation

The requirement to "adapt to changing service availability" without overwhelming Bedrock points to the AWS SDK adaptive retry mode, which layers dynamic client-side rate limiting (a token bucket that adjusts the call rate based on observed throttling) on top of standard mode's exponential backoff with jitter and circuit-breaking. Option B combines adaptive retries, exponential backoff with jitter, and a circuit breaker with a streaming handler that buffers received chunks and resumes from the last chunk, satisfying transient-error handling, streaming support, and adaptability. Option C fixes the retry mode to "standard," which does not adapt the call rate to changing availability, and its "return cached completions for failed streaming requests" would serve stale or incorrect data. Options A (fixed 1-second delay) and D (static timeouts/caps) are neither adaptive nor jittered.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!