QuestionQ41

Data Pipeline Orchestration

You need to design a data pipeline that processes large volumes of raw server log data stored in Cloud Storage. The data must be cleaned, transformed, and aggregated before it is loaded into BigQuery for analysis. The transformation requires complex data manipulation using Spark scripts developed by your team. You need to implement a solution that uses your team's existing skill set, processes data at scale, and minimizes cost. What should you do?

Explanation

Dataproc is a managed Apache Spark service that runs existing Spark scripts on scalable clusters, so it directly fits complex Spark-based transformations without requiring a rewrite. Its clusters can scale down after processing to help control costs, and Dataproc includes a BigQuery connector for loading the processed output into BigQuery.

Learn more

Community Discussion

No comments yet. Be the first to start the discussion!