QuestionQ51

Operational Efficiency and Optimization for GenAI Applications

A company is building a customer-support application that uses Amazon Bedrock foundation models (FMs) to deliver real-time AI assistance to the company's employees. The application must show AI-generated responses character by character as they are produced. It must support thousands of concurrent users with minimal latency. Responses generally take 15 to 45 seconds to complete.

Which solution meets these requirements?

Explanation

Amazon Bedrock's InvokeModelWithResponseStream returns inference output as a real-time stream, enabling partial text to be delivered as it is generated. A WebSocket-based API Gateway endpoint maintains persistent client connections so a Lambda integration can relay streamed chunks immediately, avoiding polling and complete-response buffering while supporting low-latency concurrent delivery. InvokeModelWithResponseStream — Amazon Bedrock API Reference

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!