QuestionQ26

Serving and scaling models

You have recently deployed a model to a Vertex AI endpoint and configured online serving in Vertex AI Feature Store. You have set up a daily batch-ingestion job to update your featurestore. During the batch-ingestion jobs, you find that CPU utilization is high on your featurestore’s online-serving nodes and that feature-retrieval latency is high. You need to improve online-serving performance during the daily batch ingestion. What should you do?

Explanation

Vertex AI Feature Store online-serving capacity must accommodate both online feature reads and the writes produced by batch ingestion. Configuring autoscaling for the online-serving nodes allows the store to add nodes when CPU utilization exceeds the target, reducing contention and feature-retrieval latency. Increasing ingestion workers can worsen the load, while prediction-node autoscaling addresses endpoint inference capacity rather than Feature Store serving capacity.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!