QuestionQ30

Maintaining and automating data workloads

You work for an advertising company and have developed a Spark ML model that predicts click-through rates for advertisement blocks. All development has been done in your on-premises data center, but the company is now migrating to Google Cloud. Because the data center will close soon, a rapid lift-and-shift migration is required. However, the data you have used will be migrated to BigQuery. You periodically retrain the Spark ML models, so the existing training pipelines must be moved to Google Cloud. What should you do?

Explanation

Dataproc provides managed Apache Spark and can run existing Spark ML training pipelines without rewriting the models. The Spark BigQuery connector enables those workloads to read the migrated training data directly from BigQuery, which supports a rapid lift-and-shift migration.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!