QuestionQ25

Securing AI systems

An organization has recently created a custom model that integrates with a language model (LLM). The developer observes that application programming interface (API) costs have increased. Which of the following is the best control for reducing cost?

Explanation

LLM inference is commonly priced by token usage. Adjusting token limits constrains the maximum response length, reducing generated-token consumption and therefore inference cost. AWS guidance likewise recommends controlling response length because shorter model responses lower inference cost.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!