QuestionQ37

Serving and scaling models

You have recently used XGBoost to train a model in Python that will be used for online serving. A backend service implemented in Golang, running on a Google Kubernetes Engine (GKE) cluster, will call the model prediction service. The model requires preprocessing and postprocessing steps, which must run at serving time. You want to minimize code changes and infrastructure maintenance and deploy the model to production as quickly as possible. What should you do?

Explanation

Vertex AI custom prediction routines allow a Python Predictor implementation to perform model loading, preprocessing, prediction, and postprocessing. The Vertex AI SDK builds the serving container, avoiding the need to implement and maintain a separate HTTP server or serving infrastructure, and the resulting model can be registered and deployed to a managed Vertex AI endpoint.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!