You built a custom model that carries out several memory-intensive preprocessing tasks before making a prediction. You deployed the model to a Vertex AI endpoint and confirmed that results were returned within a reasonable time. After directing user traffic to the endpoint, you find that it does not autoscale as expected when it receives multiple requests. What should you do?
Community Discussion
No comments yet. Be the first to start the discussion!
Community Discussion