QuestionQ19

Operational Efficiency and Optimization for GenAI Applications

A company is developing an API for a generative AI (GenAI) application that uses a foundation model (FM) hosted on a managed model service. The API must stream responses to lower latency, enforce token limits to control compute resource usage, and implement retry logic for model timeouts and partial responses.

Which solution meets these requirements with the LEAST operational overhead?

Explanation

API Gateway response payload streaming is supported for REST APIs and Lambda proxy integrations, allowing Lambda to relay incremental Amazon Bedrock model output rather than waiting for a complete response. Lambda is also an appropriate centralized layer for enforcing token limits and handling timeout or partial-response retries, while avoiding the operational management of a containerized inference service. API Gateway HTTP APIs do not support this response-streaming capability.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!