QuestionQ10

Serving and scaling models

You work for a textile manufacturing company. The company has hundreds of machines, each with many sensors. Your team used the sensor data to build hundreds of ML models that detect machine anomalies. The models are retrained daily, and you must deploy them cost-effectively. They must run 24/7 without downtime and provide sub-millisecond predictions.

What should you do?

Explanation

A Dataflow streaming pipeline continuously processes sensor events, and local inference with RunInference avoids the remote serving and network overhead of a Vertex AI endpoint, making it the appropriate low-latency, cost-effective pattern. RunInference can host multiple models in a Dataflow pipeline, and automatic model refresh updates the model in a running pipeline without redeploying or stopping it, supporting daily retraining and continuous availability.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!