QuestionQ17

Operational Efficiency and Optimization for GenAI Applications

A company uses Amazon Bedrock to build an AI assistant that provides customer support. Analysis indicates that 40% of customer queries use different phrasing or wording to ask the same questions.

The company needs a solution that reduces redundant model calls, ensures semantically equivalent questions receive consistent answers, and provides low latency.

Which solution meets these requirements?

Explanation

Vector embeddings represent semantic meaning, allowing differently worded but equivalent customer queries to be matched by similarity. Amazon OpenSearch Service supports vector k-nearest-neighbor search, and approximate k-NN methods such as HNSW provide faster searches on large datasets. Storing cached query-response pairs in OpenSearch and retrieving the nearest semantic match avoids redundant model calls while returning a consistent cached response.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!