QuestionQ42

Automating and orchestrating ML pipelines

You work for a pharmaceutical company headquartered in Canada. Your team has created a BigQuery ML model that predicts the number of flu infections in Canada for the following month. Weather data is released weekly, while flu infection statistics are released monthly. You must configure a model-retraining policy that minimizes cost. What should you do?

Explanation

Model retraining should be scheduled no more frequently than the availability of newly published flu infection statistics, because those monthly statistics are the outcome data needed to add newly labeled training examples. Downloading and retraining monthly avoids weekly pipeline runs and data downloads that cannot incorporate new flu outcomes, minimizing cost while keeping the model current with each new monthly data release.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!