QuestionQ2

Serving and scaling models

You built a custom model that carries out several memory-intensive preprocessing tasks before making a prediction. You deployed the model to a Vertex AI endpoint and confirmed that results were returned within a reasonable time. After directing user traffic to the endpoint, you find that it does not autoscale as expected when it receives multiple requests. What should you do?

Explanation

Vertex AI scales an endpoint based on CPU utilization by default, with a default target of 60%. Memory-intensive preprocessing may limit concurrent request handling before CPU utilization reaches that target. Reducing the CPU utilization target makes the autoscaler add replicas at a lower utilization level and therefore scale out earlier.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!