QuestionQ49

Automating and orchestrating ML pipelines

You developed a Vertex AI pipeline that trains a classification model on data stored in a large BigQuery table. The pipeline has four steps, each created by a Python function that uses the KubeFlow v2 API. The components have the following names:

Question Image

You launch the Vertex AI pipeline as follows:

Question Image

You perform many model iterations by adjusting the code and parameters of the training step. You observe high development costs, particularly for the data-export and preprocessing steps. You need to reduce model-development costs. What should you do?

Explanation

Vertex AI Pipelines reuses cached outputs only when a step’s inputs, output definition, and component specification match within the same pipeline name. Stable component names for the unchanged export and preprocessing components allow their prior outputs to be reused, while a timestamped training component ensures changed training code is rerun. This avoids repeatedly exporting the large BigQuery table and preprocessing its data.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!