QuestionQ30

Serving and scaling models

A model has been deployed on Vertex AI for real-time inference. An online prediction request returns an “Out of Memory” error. What should you do?

Explanation

An online prediction request can contain multiple instances, and processing a larger instance batch requires more serving memory. Retrying with fewer instances reduces the request’s memory requirement and can prevent an out-of-memory failure.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!