AWS Certified Generative AI Developer - Professional AIP-C01
By Amazon · Question Mode
QuestionQ19
Operational Efficiency and Optimization for GenAI Applications
A company is developing an API for a generative AI (GenAI) application that uses a foundation model (FM) hosted on a managed model service. The API must stream responses to lower latency, enforce token limits to control compute resource usage, and implement retry logic for model timeouts and partial responses.
Which solution meets these requirements with the LEAST operational overhead?
Community Discussion
No comments yet. Be the first to start the discussion!
Community Discussion