QuestionQ25

Automating and orchestrating ML pipelines

You are building a recommendation engine for an online clothing store. Historical customer-transaction data is stored in BigQuery and Cloud Storage.

You need to perform exploratory data analysis (EDA), preprocessing, and model training. You plan to rerun these EDA, preprocessing, and training steps while experimenting with different algorithm types. You want to minimize the cost and development effort required to run these steps as you experiment. How should you configure the environment?

Explanation

A default user-managed notebook VM is sufficient for iterative notebook-based EDA, preprocessing, and training without provisioning a Dataproc cluster. The %%bigquery magic lets a Jupyter notebook run BigQuery queries and convert the results to a pandas DataFrame, avoiding additional connector setup. Dataproc and Spark introduce unnecessary infrastructure cost and development overhead when distributed Spark processing is not required.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!